AI at Work

Which AI should you use to analyse business documents?

Choose an approved assistant that can access the complete evidence in a usable form and make its conclusions easy to check. For a defined bundle, compare tools using the same documents and question. For work spread across company systems, evaluate retrieval and permissions as well. No answer should be treated as complete simply because it sounds confident or includes citations.

What does the answer rest on?
  1. SuppliedWere the necessary files and appendices included?
  2. ReadDid the relevant text, tables and images reach the assistant?
  3. CurrentIs this the version that governs the decision?
  4. SupportedDoes the cited passage justify the conclusion?

A file being uploaded answers only the first question. These are review questions, not a description of a particular product's processing stages.

What happened when we tested the same records

On 8 October 2026, we gave Claude and Microsoft 365 Copilot Chat the same four short fictional records: a signed order, a proposed change, the buyer's reply and an invoice. We asked for a decision note with paragraph references. Both identified the agreed price and delivery date, and spotted the unapproved invoice charge.

The signed order was for EUR 8,000 excluding VAT, with delivery on 30 October. A later email proposed an extra EUR 1,200 and a later date. The buyer expressed interest but withheld approval. The invoice nevertheless included the extra charge and demanded immediate payment.

Our checks against the supplied records
QuestionClaudeCopilot Chat
Did it preserve the agreed price and date?Yes: EUR 8,000 and 30 October, citing A2.Yes: the same terms, citing A1 to A3.
Did it distinguish the proposal from approval?Yes. It identified the dashboard as unapproved.Yes. It said the dashboard charge lacked an approved change in the supplied records.
Did it spot the payment conflict?Yes. It compared immediate payment with written buyer acceptance.Yes. It cited both conflicting terms.
Did it invent a VAT rate?No. It said the rate and amount needed checking.No. It identified the missing rate.
What wording needed review?Its opening said no payment was currently due. The records established that acceptance evidence was absent from the bundle; they could not establish whether it existed elsewhere.Its checklist referred to approving the dashboard and revised delivery date “if those items are to be paid”. A delivery date is a contract term, not an invoice item.

What this tells a team choosing a tool

Both tools handled the central discrepancy in this small test. Their references made the conclusions easier to check, but a reviewer still needed to distinguish missing evidence from proof that something had not happened. For this task, the result supports trying the approved tool you already have and checking its reasoning against the records.

How we ran it and what it cannot establish

One fresh conversation per product, the same pasted text and instruction, with no follow-up correction. Claude displayed Opus 5.5 with Medium reasoning. Microsoft 365 Copilot Chat was used through an Acuity work account, with commercial data protection displayed and the response labelled Auto. The account interface displayed M365 Copilot (Premium); the underlying model was not disclosed. Claude's subscription tier was not recorded.

The instruction requested a note under 350 words, citations to record and paragraph IDs, and no web or connected-work search. We inspected the displayed answers against the source paragraphs. This was a short text-reasoning check. It did not test uploaded files, scans, layout extraction, long documents, SharePoint retrieval or performance across repeated runs. It does not establish a general winner or a failure rate.

All records were created for this test. No client documents were used. Full exercise materials remain part of private training delivery.

Talk to us about checking AI document analysis with your team.

Decide what analysis means for this job

Summarising a report, comparing proposals and finding an obligation are different jobs. A summary can omit detail by design. An obligation review may fail because it omitted one sentence. Asking which assistant is best at documents leaves those differences unresolved.

Define the result the reader needs. A manager comparing service proposals may need to understand exclusions and assumptions. Someone preparing a handover may need unresolved issues with an owner and a source. Judge the answer by whether it helps that reader do their job. Fluent prose can obscure a missed exclusion or an unanswered question.

This researched guidance was updated on 3 October 2026 using the official product documentation linked below.

The assistant may receive less than you can see

A person opening a file sees its layout, tables and images. An assistant's input depends on the file type and the route used to read it. Text extraction can omit an embedded image containing the crucial evidence. A scan may be legible to a person while its extracted text is incomplete or scrambled.

Anthropic's upload documentation illustrates why this needs checking: it describes visual analysis for PDFs of up to 100 pages and text-only processing for longer supported PDFs. It also says non-PDF document uploads use text extraction. A document being accepted for upload therefore does not establish that all its visual information was analysed.

Do not assume another connector or product has the same behaviour. Establish how the particular route handles the material you use. If the conclusion depends on a marked-up diagram, a handwritten note or a table spread across pages, check that content directly before relying on the answer.

A selected bundle and a company search need different checks

For a selected bundle, Claude or an appropriate Copilot experience can be evaluated on the same permitted files. Included Copilot Chat supports direct file uploads; general work-grounded access is a different capability, with licensing and configuration implications. Do not judge the included experience on information it was never supplied.

For distributed work, the retrieval route matters. Claude's Microsoft 365 connector can search supported organisational sources with user-delegated access. Copilot has its own work-grounded experiences. Neither route gives a search result authority simply because it is accessible. An old attachment can still be the wrong version.

Selected document bundleSearch across company systems
Check that every required file is suppliedCheck that the relevant source can be found
Identify the current version before analysisEstablish which retrieved version has authority
Check extraction and page referencesCheck retrieval coverage and the opened source
Record missing appendices or unreadable pagesRecord repositories or permissions outside the search

A useful answer makes disagreements visible

Imagine two fictional maintenance proposals. One price includes emergency visits; the other lists them separately. A summary that puts both headline prices in a table may make the cheaper proposal look preferable without showing that the scope differs.

The useful analysis identifies the difference, points to the relevant passages and leaves the commercial decision with the reader. It should distinguish a stated exclusion from something the document simply does not explain. 'Not found' and 'not included' are different conclusions.

Source references are a way into the evidence. Open them. A reference to the right document can still point to a passage that does not support the claim. For a material conclusion, the reviewer needs to see the relevant wording in context, including any qualifications nearby.

Compare the difficult cases that affect your work

A clean, short document is a useful starting point, but it says little about performance on conflicting versions or scanned attachments. Include representative difficulty in an evaluation using permitted or independently created material. Keep the same evidence and task across products so you can explain any difference.

Look particularly at omissions. A plausible incorrect answer is visible once someone checks it; a missing issue may never reach the reviewer. An experienced person should identify the important issues independently before using the AI output to assess coverage. That gives the comparison a standard beyond how convincing each response appears.

Include the effort required to prepare files and resolve uncertainties. A tool that produces a faster initial summary can still consume more time if the reviewer repeatedly has to locate evidence or repair page references. Record those difficulties as part of the result.

Choose a workable review process before expanding access

Start with the approved tool your team already has when it can handle the material. Consider an alternative when a specific limitation persists and a fair comparison demonstrates a useful improvement. Account approval, retention and permitted use of the documents need to be settled before staff move information into a new service.

Different jobs may justify different choices. A defined research bundle can favour a carefully prepared chat workflow. Repeated enquiries across current internal sources may favour an integrated retrieval route. Highly consequential interpretation still needs the appropriately qualified person to own the conclusion.

For an organisation adopting AI document analysis, the practical investment includes agreeing which sources count, who reviews important conclusions and where the reviewed output is kept. Keep that process with the team doing the work so a change of assistant does not quietly change the standard of review.

Sources and product references

Checked 3 October 2026. Features and access can change, so check what your own licence includes.

  1. Anthropic: document uploads, extraction and PDF processing
  2. Microsoft: Copilot Chat file uploads and work-data access
  3. Anthropic: Microsoft 365 connector setup and delegated permissions

Editorial guidance informed by Acuity's work with teams. Examples are fictional and client materials stay private. To record who owns an AI use and when it is reviewed, see AI Register.