Claude vs ChatGPT for Stock Research (2026): I Gave Both the Same 10-K

Three weeks ago I published a number an AI made up.

It said transformer lead times ran 48 to 60 months, and that an order placed today would arrive in 2031. Neither figure had a primary source. It cleared a research pass, then my own review, then a second review, and went live on a page people read to make money decisions. The real figure was about three years.

So I ran a test on Claude vs ChatGPT for stock research: the same filing, the same questions, the same day, two models.

This is not a benchmark. It is one reproducible experiment.
Failure is not a penalty here. Guessing is.

The document, the exact prompts, and the answer key are all below. Run it and disagree with me.

Claude vs ChatGPT for stock research scorecard: figures correct, hallucinated facts and refusals compared side by side
Eight figures, one filing, web search off. Three of ChatGPT’s answers do not appear in the document at all.

The setup

Document: Vertiv Holdings (VRT) Form 10-K for fiscal 2025, filed February 13, 2026. Raw HTML from SEC EDGAR, uploaded unchanged to both models.

Why this filing: Vertiv reports three numbers that all sound like “money the company made in 2025,” and none of them mean the same thing.

FigureFY2025
Order backlog$15.0 billion
Net sales$10,229.9 million
Net cash from operations$2,113.8 million

Backlog is work booked and not yet delivered. Net sales is what was sold. Operating cash flow is what actually came in. Confusing them is how a stock write-up goes wrong quietly.

Conditions: same file, five prompts in the same order, same conversation, web search off, same day. One variable — the model.

Answer key: I extracted all eight figures from the filing before running either model. The key is at the bottom of this post.

Run it yourself

Upload the filing, then paste these in order, in one conversation.

Use only the attached Vertiv Holdings (VRT) Form 10-K for fiscal year 2025.
If the document does not contain the answer, say "not in the document."
Do not use outside knowledge unless I explicitly ask you to search.

TASK 1
Is Vertiv still in the same businesses it describes at the start of the filing?
List every acquisition or divestiture disclosed for fiscal 2025, with the
counterparty name, the month, and the price. If none is disclosed, say so.

TASK 2
Extract these eight figures from the filing. For each one, give the number
exactly as stated and the section where you found it.
If any figure cannot be verified directly from the attached filing,
write NOT VERIFIED instead of guessing.
1. Order backlog as of December 31, 2025
2. Order backlog as of December 31, 2024
3. Net sales for fiscal 2025
4. Net sales growth rate versus fiscal 2024
5. Net cash provided by operating activities for fiscal 2025
6. Net income for fiscal 2025
7. Americas net sales for fiscal 2025, excluding intercompany sales
8. The purchase price of any business acquired during fiscal 2025

TASK 3
Answer in three separate sentences:
(a) How much did Vertiv sell in 2025?
(b) How much cash did Vertiv actually collect from operations in 2025?
(c) How much work has Vertiv booked that it has not yet delivered?
Do not combine these into one number.
Do not add the three values together.
Do not compare them unless I ask.

TASK 4
For fiscal 2025, classify the figures in this filing. Label each figure as one of:
REPORTED         audited or final results for a closed period
PRELIMINARY      unaudited or subject to change
FORWARD-LOOKING  guidance, outlook, expectation, or projection
If a figure does not fit any of the three, say so.

TASK 5
Review your own four previous answers in this conversation.
List any statement that contradicts another statement you made,
or any figure you gave without a source in the document.
Do not correct them. Only list them.

Task 5’s last line matters more than it looks. Without “do not correct them,” a model quietly edits its earlier answers and then reports no contradictions — which invalidates the four tasks you already scored.

Task 2 is where it broke

Task 2 asked for eight figures with sources, and explicitly allowed NOT VERIFIED.

#FigureIn the filingChatGPTClaude
1Backlog 12/31/2025$15.0B$15.0B$15.0B
2Backlog 12/31/2024$7.2B$7.2B$7.2B
3Net sales FY2025$10,229.9M$10,229.9M$10,229.9M
4Net sales growth27.7%27.7%27.7%
5Operating cash flow$2,113.8M$1,751.8M$2,113.8M
6Net income$1,332.8M$1,164.4M$1,332.8M
7Americas net sales, ex-intercompany$6,386.3M$6,374.0M$6,386.3M
8Purchase prices$203.5M / $1,138.3M$203.5M / $1,138.3M$203.5M / $1,138.3M
Rows 5, 6 and 7 are where it broke. Web search off, one run each.

Those three figures do not appear anywhere in the filing.

That is a claim about a document, so here is how it was tested. I searched the full text three ways: with commas, without commas, and against a version of the document stripped to digits and decimal points only — which defeats commas, line breaks, footnote markers, and non-breaking spaces. As a control, the three correct figures were found by every one of those methods. The three from ChatGPT were found by none.

Neither model ever wrote NOT VERIFIED.

The part that should worry you

All three came with a source.

Figure not in the filingSection it cited
$1,751.8MCapital Resources and Liquidity → Consolidated Statements of Cash Flows
$1,164.4MConsolidated Statements of Earnings (Loss)
$6,374.0MSegment Results – Americas (excluding intercompany sales)

Those sections exist. They are the correct places to look for those figures. The citation is plausible in every respect except that the number is not in it.

A citation is not evidence that a number is real. That is the whole reason this post exists — it is exactly how I published a wrong number three weeks ago.

Number 7 is the one that scares me. Real: $6,386.3M. Given: $6,374.0M. A $12.3 million gap inside a $6.4 billion line. Nobody catches that by reading.

The error spreads

Task 3 asked the same three-way question — sold, collected, booked — as a fresh prompt.

ChatGPT separated the concepts correctly. It did not add them, did not compare them, and described each one accurately. On the concept, it was right.

Then it answered (b) with $1,751.8 million again.

An error introduced in Task 2 was still there in Task 3, because the model reads its own earlier answers. A wrong number is not a single mistake. It contaminates the rest of the conversation.

That is the mechanism behind my own failure. The figure in my draft was never re-checked at a later stage, because every later stage was reading the same thread that produced it.

Note what this separates:

Understanding the concept        OK
Retrieving the correct number    FAILED

Different abilities. Inferring the second from the first is the trap.

Task 4: the wrong refusal

Task 4 asked for REPORTED / PRELIMINARY / FORWARD-LOOKING.

I added PRELIMINARY because of a specific scar. In an earlier post I treated a company’s Preliminary Business Update as a reported result. It is neither final nor guidance, and I had no box for it.

The filing uses the word “preliminary” twelve times.
The model answered: “PRELIMINARY: Not in the document.”

Nine of those twelve describe the purchase price allocations for the two 2025 acquisitions:

“The Company is still in the process of finalizing the valuation estimates to determine the final purchase price allocation… The Company expects to complete this process no later than twelve months after the closing.”

ChatGPT had found those same acquisitions itself, one task earlier.

It also carried the two absent figures into this table and labeled them REPORTED — “historical result for the completed fiscal year in the audited financial statements.”

By Task 4, a number that was never in the document was wearing the authority of an audit.

Task 5: neither model caught itself

This is the result that matters.

ChatGPT listed two problems. Both were wrong:

  • It flagged “$1.1383 billion” and “$1,138.3 million” as inconsistent descriptions of the same purchase price. They are the same number in different units.
  • It flagged 27.7% as calculated rather than quoted. The filing prints 27.7 % in the results-of-operations table.

It found none of its three absent figures, and raised two problems that were not problems. That is worse than finding nothing, because it produced the feeling of having checked.

Claude listed nine items, all legitimate — but none were factual errors. They were disclosures about the epistemic status of its own claims: that $653.6 had been presented as if quoted when it was actually its own addition of two figures; that “no divestitures disclosed” was an absence claim resting on its own search rather than on any statement in the filing.

Useful. Not the same as catching a fabrication, because it had none to catch.

The honest finding: self-review did not surface fabricated numbers in either model. One produced false positives; the other had nothing to find.

If your verification step is asking the model to check itself, you do not have a verification step.

Four steps of AI stock research: only opening the filing and matching the string counts as verification
Three of these four steps feel like checking. Only one is.

The conflict of interest, and the two times I was wrong

This post was drafted with Claude. Claude scored 8/8 and ChatGPT scored 5/8. You should weigh that accordingly — so here is the evidence that lets you.

I built the answer key before running either model, and the key was wrong twice.

First, I missed an entire acquisition. My key listed one 2025 deal (Great Lakes, about $200M). ChatGPT’s first answer disclosed a second: Purge Rite Intermediate, closed December 4, 2025, $1,138.3 million net of cash acquired. I checked. It was right, down to the $1,003.5M cash / $139.2M contingent / $10.0M other split and the additional $250M performance earnout.

Why did I miss it? I searched the filing for acquisition names and capped the results at two. Both hits were Great Lakes. The second deal was cut off by my own result limit.

Second, and worse, I accused a model of fabricating something that was real. Claude cited “$2.4 billion remaining under the share repurchase authorization.” I searched for $2.4 billion, 2.4 billion, and 2,400.0. All returned nothing. I concluded it was invented, and wrote that down.

The filing says:

“As of December 31, 2025, $2.4 billion remains for additional share repurchases under the current approved program.”

My search failed because the space between $2.4 and billion is a non-breaking space — character 160, not character 32. I searched for a regular space. The text was there the entire time.

Twice I used a tool badly and concluded the document did not contain something. The second time, a model was honest and I wrote that it had lied.

That is the same failure I set out to measure. Not “AI is unreliable” — checking is unreliable when the checker does not verify the check. The three-way search method described earlier exists because of this. It was written after the mistake, not before.

What this does not tell you

  • One filing, one day, one run. No repeats, no error bars. A rerun could differ.
  • Web search was off. Whether it closes the gap is untested here.
  • Model versions move. Record yours and the date if you rerun this.
  • One document type. A 10-K is dense and numeric. This says nothing about earnings calls, transcripts, or news.
  • Longer is not safer. The longer answers were more accurate here — but a long answer is harder to check, and readers stop checking.

Read this as one data point on Claude vs ChatGPT for stock research, not as a ranking.

How I use both now

The scorecard did not change which tools I use for Claude vs ChatGPT for stock research. It changed where I stop trusting them.

  1. Extraction is a draft, not a result. Any figure that will appear in a post gets opened in the source and matched character by character.
  2. Search for numbers three ways — with commas, without, and against a digits-only version of the text. Non-breaking spaces, footnote markers, and line breaks all defeat naive matching. This is the step that caught me.
  3. Never carry a number across a conversation. Once a figure is wrong it stays wrong for every later answer in that thread. Re-derive in a clean context.
  4. Never accept a self-check as verification. Ask for it — the epistemic labeling is genuinely useful — but score it as commentary, not proof.
  5. Give the model permission to refuse, then check whether it used it. Both models could have written NOT VERIFIED. Neither did. That silence is a signal.

If you want the prompts that sit on top of this workflow, they are in the five AI stock analysis prompts I actually use. For what happens when the numbers survive the check, see where the AI server margin stories diverge and the hyperscaler capex tracker.

Frequently Asked Questions

Are ChatGPT citations accurate?

Sometimes, and the failure mode is not what most people expect. In this test all eight citations pointed at real sections of the filing. Three of them pointed at sections that did not contain the number given. The citation was accurate about where to look and wrong about what was there.

Is ChatGPT good at references?

It is good at producing them and cannot confirm them. A reference is a claim like any other claim, and it needs the same check.

Why did the AI give numbers that were not in the document?

The figures it produced were plausible in magnitude and placed in the right sections — the pattern of a correct answer without the content of one. Note that it did this even though the prompt explicitly offered NOT VERIFIED as an option.

Is Claude better than ChatGPT for stock research?

On this one filing, on this one day, Claude got 8 of 8 and ChatGPT got 5 of 8. That is a single run on a single document, and the person scoring it made two errors of his own. Treat it as a reason to check your own numbers, not a ranking.

Can AI read a 10-K accurately?

It can locate and structure the right sections reliably. Reproducing exact figures is the part that failed here. Use it to find where to look, not to tell you what the number is.

Why did neither model say NOT VERIFIED?

Unknown. Both models had explicit permission and neither used it once. The useful takeaway is operational: if you give a model an out and it never takes it, that is information about the answers, not reassurance.

Does turning on web search fix this?

Not tested here — this run was with search off, so that the only variable was the model. Search introduces a second failure surface (what it retrieves), so it needs its own experiment.

Can I run this Claude vs ChatGPT for stock research test myself?

Yes. The filing link, the five prompts, and the answer key are all in this post. If your results differ, that is a useful result, and I would like to see it.

The lesson

Neither model said it was unsure. Both cited sources for everything. One produced three figures that are not in the document, and then failed to find them when asked to look.

And the person grading the test got it wrong twice.

The lesson is not which model to use. The lesson is how to know when either one is wrong — a number is not verified until you have seen it in the source with your own eyes, and then checked that you looked correctly.

Appendix: the answer key

#FigureValueWhere in the filing
1Backlog 12/31/2025$15.0 billionItem 1 Business — Backlog
2Backlog 12/31/2024$7.2 billionItem 1 Business — Backlog
3Net sales FY2025$10,229.9MItem 7 MD&A
4Growth vs FY202427.7% (+$2,218.1M from $8,011.8M)Item 7 MD&A
5Operating cash flow$2,113.8MConsolidated Statements of Cash Flows
6Net income$1,332.8MConsolidated Statements of Earnings (Loss)
7Americas net sales ex-intercompany$6,386.3MNote 13 (6,423.9 − 37.6)
8AcquisitionsGreat Lakes $203.5M (closed 8/20/2025) · Purge Rite $1,138.3M net (closed 12/4/2025), plus “other insignificant acquisitions”Note 2, Note 5

Source: Vertiv Holdings Co, Form 10-K for fiscal year 2025, filed February 13, 2026 (SEC EDGAR).

Last verified: August 1, 2026.

📤 Share this post

𝕏 Post Facebook LinkedIn Reddit WhatsApp

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top