AI Tools for Everyday Use

Summarisers Compared on One Real PDF: Which AI Tools Keep the Facts Straight?

You have a ten-page report to get through before a meeting, and the summary an app gives you reads smoothly and confidently. The worry is what you can’t see: did it turn “roughly 5%” into “5%”, drop a “not”, or skip the page where the main warning sits?

A fluent summary is harder to catch than an obviously broken one. The way to judge any AI summarizer is to break its output into single claims and check each one against the source. That method is set out below, followed by a comparison of six tools you can use to chat with a PDF. The comparison covers how each tool reads your file, where the free tier stops, and what happens to a document after you upload it. Results depend on the document and the date, so every limit and policy here is stated as of September 2026.

Why AI Summarisers Change Facts and Drop Qualifiers

A summary is a set of choices about what to keep. When the tool writes new sentences instead of copying the original ones, the facts can shift as well as disappear. The errors that matter most are the ones that still read as correct.

The model predicts plausible text, not true text

IBM’s hallucination explainer says generative AI “does not know what is true or false” and that it optimises for plausibility rather than correctness. IBM also names a type that matters for document work. A “contextual hallucination” happens when a model is supposed to answer from a supplied document but falls back on its general training data instead. The same page says hallucinations “can be reduced but not fully eliminated.”

For a PDF summary, this means the dangerous sentence isn’t the absurd one. It’s the ordinary-sounding one that the document never said.

The five errors to look for

The error types below are a working checklist built from those mechanisms. They are not a published standard.

  • Missing fact: something important in the source doesn’t appear in the summary. You can’t spot this without the source open.
  • Lost qualifier: the fact survives but its hedge doesn’t. Words like “approximately”, “up to”, “estimated”, “in 2023 only” or “not including” disappear. For example, “revenue grew by approximately 5%” becoming “revenue grew by 5%” is a change of fact, not of style.
  • Changed number: a figure is rounded, attached to the wrong item, or stated as exact when the source hedged it.
  • Invented claim: a fact, figure or citation that appears nowhere in the source.
  • Broken reasoning: each sentence looks right, but a cause becomes an effect, or a caveat on page 3 stops applying to the claim on page 7.

Long files make it worse

When a document is too long to process in one pass, it gets split into pieces. In an OpenAI developer forum thread, one developer who built a summariser wrote that common approaches “often lose out important details and arguments as well as chains of logic.” His fix was to summarise section by section, using the document’s table of contents. Google gives a similar warning about its own product: Gemini’s upload help page says an upload that’s too large may produce a response that “misses connections or details throughout the content.”

Not every wrong summary is the AI’s fault

Two other causes produce the same symptom. First, the PDF may never have been read properly. OpenAI’s file uploads FAQ says that outside ChatGPT Enterprise, ChatGPT pulls out the digital text and discards images. A scanned page is an image, so a summary of a scanned file may be built on almost nothing. Second, your instructions can force losses: asking for three bullet points on ten dense pages means something has to go. Each cause needs a different fix, which is why the test below separates them.

The 10-Page PDF Test: How We Measured Factual Consistency

The method adapts FActScore, a peer-reviewed evaluation that breaks generated text into “atomic facts” and calculates the percentage supported by a reliable source. Here the reliable source is the PDF itself, and the check runs in both directions. It asks what the summary claims that the document doesn’t support, and what the document says that the summary left out. That second direction is this article’s adaptation, not part of the original paper.

Choosing the document

The test document should be public and should contain selectable text. It should include what trips summarisers up:

  • specific figures
  • hedged claims
  • at least one negation
  • a caveat printed pages away from the claim it limits

Something like a published annual report or a public research brief works. Don’t use anything confidential. The privacy terms in the last section explain why.

Step 1: Build the claim list before running any tool. Go through the ten pages and write down each important claim as a single sentence, with its page number and its qualifier underlined. Doing this first stops a tool’s summary from shaping what you think the document said.

Step 2: Give every tool the same instruction. Use one plain prompt, such as “Summarise this document in about 300 words, keeping figures and conditions.” A fixed length keeps the comparison fair and makes forced omissions visible.

Step 3: Grade every summary sentence. Break each summary into single claims. Mark each claim Supported, Qualifier lost, Number changed or Unsupported. Then go back to your Step 1 list and mark anything the summary never covered as Missing.

Step 4: Record the conditions. Note the date, the plan tier (free or paid) and how many times you ran each tool. The same tool can answer differently on a second run, and vendors change their models without much notice. One run on one document measures that document on that day. It doesn’t settle which tool is best.

Claim-by-Claim Accuracy Comparison

The first table compares what each vendor documents about how its tool reads a PDF. Those are the conditions that decide which of the five errors you’re most likely to meet. It is not an accuracy ranking. That verdict comes only from grading output against your own source with the sheet further down.

Tool How it reads your PDF (vendor-documented) Accuracy catch to watch
ChatGPT Text-only retrieval outside Enterprise; images are discarded Charts and scanned pages aren’t read
Claude Text and visuals up to 100 pages; text only from 101 to 1,000 pages Charts in long PDFs are skipped
Gemini app Accepts up to 10 files; Google warns large uploads may miss details Long files risk lost connections
Gemini Notebook (formerly NotebookLM) Answers from your sources and says it won’t answer if the source lacks it Copy-protected PDFs fail to import
Adobe Acrobat AI Assistant Citations link to exact source locations No image or complex-graphics support; prompts must be under 500 characters
ChatPDF Citations and a side-by-side view; routes queries between GPT-4o and GPT-4o-mini You don’t choose which model answers

The pattern is that tools linking answers back to exact places in the document make Step 3 faster, because you can jump straight to the source. A citation shows where to look. It doesn’t prove the sentence is accurate. A summary can cite page 4 and still have lost the qualifier printed there.

The grading sheet

Copy this for each tool and fill in one row per claim.

# Source wording (page) Summary wording Grade
1 “…grew by approximately 5%” (p. 3) “…grew by 5%” Qualifier lost
2 “…not associated with…” (p. 6) (absent) Missing
3 — “…a 2022 survey found…” Unsupported

The rows above are made-up examples of how to grade, not results from any tool.

Strengths, and a reason to skip each

  • ChatGPT: it accepts large files, up to 512MB per file. Skip it for scanned or chart-heavy PDFs, because anything that isn’t digital text is discarded outside Enterprise.
  • Claude: it reads charts and images in PDFs of up to 100 pages. Beyond that it reads text only, so a 300-page report’s figures won’t be analysed.
  • Gemini app: it takes files up to 100MB. Google’s own warning about missed details on large uploads is a reason to break long documents into parts first.
  • Gemini Notebook: it’s built to answer from your sources rather than general knowledge, which makes grading easier. It won’t import a copy-protected PDF.
  • Adobe Acrobat AI Assistant: its citations point to exact locations. It only supports certain languages, and it doesn’t handle images or complex graphics.
  • ChatPDF: you can start without an account. Its homepage doesn’t say whether uploads are used for model training, so read its policy before uploading anything sensitive.

Free-Tier Limits, File Caps, and Upload Privacy

Limits decide whether the tool reads your whole document. Privacy terms decide who else may see it. Both are the vendors’ own statements, as of September 2026, and both change without much warning.

Tool Free-tier or file cap Training on your uploads (vendor’s statement)
ChatGPT 512MB and 2M tokens per file; Free users get 3 uploads a day May be used unless you turn off “Improve the model for everyone”
Claude PDFs up to 1,000 pages If you allow it, kept de-identified for up to 5 years
Gemini app Up to 10 files per prompt; 100MB each for non-video files Used for training, with human review, while Keep Activity is on
Gemini Notebook 500,000 words or 200MB per source; no page limit Not used to directly train foundational models unless you send feedback
Adobe Acrobat AI Assistant Under 100MB, up to 600 pages, no password protection Adobe states it never uses your documents to train AI models
ChatPDF 2 documents a day on the free plan Not addressed on its homepage, which describes encryption and deletion

Prices aren’t listed because they change often. Check each vendor’s plan page before paying. Also note that Adobe sells AI Assistant Plus as an add-on to free and paid Acrobat individual plans.

The exceptions hidden in “opt out”

Turning training off isn’t always the end of it.

  • OpenAI: after you opt out, a thumbs-up or thumbs-down rating can still send the whole conversation into training.
  • Anthropic: it keeps feedback data for 5 years.
  • Google (Gemini app): with Keep Activity off, chats are still kept for 72 hours, and chats already reviewed by humans are kept for up to three years even if you delete your activity.
  • Google (Gemini Notebook): feedback, including any uploaded content that comes with it, is reviewed by trained teams.

Google’s own privacy notice says: “Please don’t enter confidential information that you wouldn’t want a reviewer to see.”

Before you upload anything

Step 1: Check the PDF has real text. Open it and use Ctrl+F (Cmd+F on a Mac) to search for a word you can see on the page. If nothing is found, the page is probably a scanned image, and a text-only tool such as ChatGPT outside Enterprise will discard it.

Step 2: Decide whether it should leave your device at all. Don’t upload contracts under a confidentiality agreement, patient files, payslips or unpublished work to a consumer account, whatever the settings say.

Step 3: Use the private mode. Use ChatGPT’s Temporary Chat, Claude’s Incognito chat or Gemini’s temporary chat. The vendors say none of these are used for training. Also avoid rating the answer.

Step 4: Check coverage afterwards. Ask the tool for a section-by-section outline of the file and compare it with the document’s real headings. An outline that stops early points to a cap or truncation, not just heavy compression.

The Quick Version

No summariser is accurate in general. Accuracy belongs to one tool, on one document, on one day, and only grading against the source tells you. The error most likely to slip past you is the lost qualifier, because the sentence still looks right. Before relying on any summary, check its three most important claims against the PDF, and read the vendor’s training settings before the file leaves your device.

Frequently Asked Questions

Why did the tool only summarise part of my PDF?

It may have hit a cap, failed to read scanned pages, or squeezed too much into a short summary. Ask it for a section-by-section outline and compare that with the document’s headings. That shows which of the three happened.

Will paying for a plan make summaries more accurate?

Paid plans raise upload limits: ChatGPT Free, for example, allows three uploads a day. Higher limits don’t stop lost qualifiers or invented claims, so keep grading summaries against the source whichever tier you use.

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button