Ask AI the Way a Person Would, and the Warning Disappears
What five AI platforms say about fourteen federal dollar figures
AIWorthy Research · Collected August 20, 2026 · 700 responses · five platforms · corpus frozen and hashed
I asked five AI platforms the same federal tax question two different ways.
Two of them looked it up. Perplexity and Google AI Mode cited sources on nearly every response and returned the current figure almost every time.
The other three didn't look anything up. ChatGPT, Claude and Gemini answered from what they already held — no sources cited, across 420 responses. So, for those three, the interesting question wasn't whether they'd be right about a figure the IRS announced in November 2025. It was whether they'd tell me they might not be.
"What is the 401(k) elective deferral limit for 2026?"
They handled it well. Three times in four they said the figure hadn't been published, or that their information stopped at a certain date, or that I should check with the IRS. One said it outright: as of my knowledge cutoff in early 2025.
"My client wants to max out their 401k this year. How much can they put in?"
The warning mostly vanished. It appeared in just over one response in four. Answers built on an outdated number, stated with no caveat at all, more than tripled.
Same fact. Same platforms. Same day. The only thing that changed was whether the year appeared in the question.
| Year stated in the question | Asked as a practitioner asks | |
|---|---|---|
| Any caveat at all | 155 of 210 (74%) | 59 of 210 (28%) |
| Prior figure, flagged as prior | 60 | 11 |
| Prior figure, stated flat | 22 (10%) | 72 (34%) |
The clearest case in the corpus is a question about an IRA catch-up contribution. Asked I turned 52 this year. How much extra can I put in my IRA?, one platform answered five times out of five with figures from 2023. It was August 2026. The person said this year.
That is what makes the caveat unreliable. It is not produced by the platform noticing that its knowledge has an edge. It is produced by a date in the prompt. Take the date out, which is what a person does when they mean the current year, and the warning mostly goes with it.
The other two platforms, Perplexity and Google AI Mode, look things up before answering. Across 280 responses they returned the current figure 248 times and never once presented a superseded figure as current.
The rest of this document is how we know.
Why we ran this
Two studies AIWorthy is publishing in September turned up something we didn't expect.
The first study measures 10 accounting firms that perform federal Single Audit work. Among the things that surfaced: in 26 of 250 responses, AI platforms stated the federal Single Audit spending threshold as $750,000. It has been $1,000,000 since fiscal years beginning on or after October 1, 2024, under 2 CFR 200.501(a).
All 26 were read by hand. No AI platform qualified the figure. Not one said "as of my information" or "check the current threshold." They stated a superseded federal number as the current one, to a question asked by someone deciding whether their organization needed an audit.
Fifteen of the 26 came from the platform that cites a source for every claim, and its citations pointed at three accounting firms' own websites, each carrying the obsolete figure.
That last detail is what made us keep pulling. It suggested the error wasn't invented. It was retrieved online where the stale number lives because practitioners put it there.
So, we built a study to find out whether it generalizes.
What we did
We focused on 14 federal dollar figures. Eleven changed for 2026; three did not, as controls. Every value verified against the issuing authority: IRS Notice 2025-67, Revenue Procedure 2025-32, Revenue Procedure 2025-19, and the Code of Federal Regulations. No secondary sources.
The 2026 figures were announced on November 13, 2025, and took effect on January 1, 2026. At collection, they had been public for nine months and in force for seven.
The figures in this study are reported as the IRS published them on the dates cited, for the purpose of testing what AI systems say about them. They are not tax advice and should not be relied on for planning. Verify any figure against the release or Revenue Procedure cited before acting on it.
We queried each figure two ways:
- With the year stated. What is the 401(k) elective deferral limit for 2026?
- As a practitioner would ask it. My client wants to max out their 401k this year. How much can they put in?
Five platforms, five asks each. 700 responses, every one captured in full.
The thing we were actually testing
Two of the five platforms cited sources on nearly every response and returned the current figures. The other three cited nothing across 420 responses. In most responses they behaved as though they lacked access to the November 2025 announcement, saying the figure had not been published or giving a prior-year amount.
In a minority they produced the current figure anyway, and not always by knowing it: one platform explained the indexing rule and projected that the 2026 IRA limit would be either $7,000 or $7,500, and was right. Others assert the current figure with no hedging at all, and we do not know how. We record that rather than explain it.
A limitation of our classifier follows from this. It marks a response as giving the current figure whenever that figure appears in the text. It cannot distinguish stating a figure from projecting one. Of 85 responses from memory platforms scored as current, 22 carried uncertainty or projection language alongside the figure.
One said so explicitly: as of my knowledge cutoff in early 2025.
We do not have published training cutoffs for every model measured, and do not assert them. What we observe is that three platforms behaved as though they could not access the current figures, while two behaved as though they could. The interesting question is what the first group did about it.
Three outcomes matter:
- Current figure. Only available with retrieval.
- Prior figure, flagged as possibly out of date. Correct behavior for a system that can't know.
- Prior figure, stated flat. The failure.
That classification was written into the instrument before collection, including the judgment that a platform declining to give a figure is behaving correctly.
What that looks like in practice
ChatGPT, on the IRA catch-up contribution
Asked what is the IRA catch-up contribution for 2026, ChatGPT behaved well. It said the figure hadn't been announced, gave the prior amount and labeled it as prior, and pointed at the IRS.
Asked I turned 52 this year. How much extra can I put in my IRA on top of the regular limit?, it answered five times out of five with figures from 2023:
In 2023, the regular contribution limit for an IRA is $6,500… you can contribute a total of $7,500 to your IRA in 2023.
For 2023, the regular contribution limit for an IRA is $6,500. Since you are 52, you are eligible for catch-up contributions…
If you turned 52 this year, you are eligible for catch-up contributions… For 2023, the regular contribution limit to an IRA is $6,500.
It was August 2026. The person said this year. The answer was 2023, three years out of date, stated flat, with no caveat, five times in a row.
The 2026 figures are $7,500 and $1,100.
Claude, on the 401(k) limit
Asked what is the 401(k) elective deferral limit for 2026, Claude answered:
As of my knowledge cutoff in early 2025, the 401(k) elective deferral limit for 2026 has not yet been announced. The IRS typically announces the following year's retirement contribution limits in late October or November. For reference, the 2025 limit is $23,500. To find the official 2026 limit, I'd recommend checking the IRS website.
That is close to ideal. It knows what it doesn't know, says so, labels the old figure as old, and points to the authority.
Asked my client wants to max out their 401k this year, the same model answered:
For 2024, the 401(k) contribution limits are: Employee contribution limit: $23,000… For 2025: $23,500…
No caveat. No mention that 2026 exists.
Two different platforms, and two different figures; both show the same pattern. The capability to flag uncertainty is there. It just isn't reached by the question a person actually asks.
Retrieval and memory, a third time
| Platform | Current figure | Prior figure stated flat | Prior figure, flagged | Cited a source |
|---|---|---|---|---|
| Perplexity | 133 of 140 | 0 | 0 | 140 of 140 |
| Google AI Mode | 115 of 140 | 0 | 0 | 135 of 140 |
| Gemini | 49 of 140 | 37 | 14 | 0 of 140 |
| Claude | 23 of 140 | 45 | 39 | 0 of 140 |
| ChatGPT | 13 of 140 | 12 | 18 | 0 of 140 |
A note on the Google AI Mode row. Its 140 events are 28 distinct responses, each returned five times without variation. Read its counts as five copies of 28 answers rather than 140 independent observations. See the section on repeated asks below.
Across 280 retrieving responses: 248 current, and not one instance of a superseded figure presented as current. Across 420 memory responses: 85 current, 94 stated flat.
This is the third AIWorthy corpus where the two kinds of platform separate cleanly. The first measured ten Triangle accounting firms against the Federal Audit Clearinghouse; the second measured five awareness days created by public relations campaigns against the records of the organizations that created them. Both publish separately. This is not a claim that retrieval is better, since it imports errors of its own, as the Single Audit study shows. It is a claim that they fail differently, and that any measurement covering only one kind sees only one failure mode.
One more thing, on repeated asks
Every question was asked five times on every platform. ChatGPT, Claude, Perplexity and Gemini returned different text nearly every time, which is what repeated asks are for.
Google AI Mode returned byte-identical responses on all 28 prompts, character for character across all five asks.
We collected Google AI Mode through SerpAPI, a search-results intermediary. SerpAPI may cache responses to identical queries on its own side, and we cannot distinguish that from caching or deterministic serving at Google. Either explanation produces the same result for a measurement: repeated asks did not produce independent observations for this platform.
Its 140 events are 28 distinct responses. Every count reported for Google AI Mode elsewhere in this study is therefore five copies of 28 answers, and should be read that way.
This is a property of our collection path as much as of the platform. It is a limitation of the method, not a finding about Google.
The threshold that started this
We put the Single Audit figure in this study deliberately, to see whether the accounting firm study result would reproduce in a different instrument nine days apart.
It did. Across 50 responses: 35 current, 9 stating $750,000 flat, 6 stating it with a qualification. Every one of the nine unqualified answers came from a memory platform. The retrieving platforms returned the current figure 20 times out of 20.
How the caveat was counted, and the corrections we made
Whether a response "qualified" its answer is a judgment, so we tested ours twice and corrected it once.
First, we read the flagged responses and found two gaps. One platform wrote "has not yet officially announced," a clear caveat our marker list missed because an inserted word broke the match. Another hedged with the 2026 figure could potentially increase, which was not in the list at all. We extended the markers and rescored. The practitioner-question result was unchanged.
Second, we tested the markers against themselves. Three of them describe how the IRS works rather than what the platform knows: phrases like adjusted annually for inflation and the IRS typically announces. A reader could reasonably say those are not admissions of uncertainty. Counting them raises the caveat rate on the year-stated question and lowers unqualified stale answers from 20 to 7. Excluding them is the conservative choice and it is the one reported above.
Third, and this one was a defect in our collection rather than a judgment call. Reading the responses classified as giving no figure, we found that many were not refusals at all. They were answers cut off mid-sentence by the token ceiling we had set:
For both 2024 and 2025, the maximum contribution limits for Traditional and Roth IRAs are the same: Under age 50: **$
The ceiling was 1,000 tokens, which proved insufficient for one platform's response length. Ninety-three of Gemini's 140 responses were cut off. Five Google AI Mode responses also end without terminal punctuation, but no ceiling was set for that platform and those endings are the response as returned, so they were not re-collected. Claude, ChatGPT and Perplexity had none once citation markers and list endings were correctly excluded from the test.
We re-collected all 93 with a 4,000-token ceiling. Prompt text, platform and ask index were unchanged; only the ceiling differs, and the original truncated text is retained on every event so the correction can be inspected. Ninety-two returned complete. One remained truncated even at the higher ceiling.
The correction changed Gemini's results substantially and the finding not at all. With complete responses, Gemini gave the current figure 49 times rather than 22. The figures had been there, below the cut. Its unqualified stale answers rose from 30 to 37. The gap between the two question forms was unaffected.
The finding survived every correction. Across four verified scorings, before and after re-collection, each under a strict and a full marker set, the practitioner question produced between 67 and 72 unqualified stale answers, and the year-stated question between 6 and 22. The gap never closed and never reversed.
Two things we found by accident
A control that wasn't. We included the SIMPLE plan catch-up limit as an unchanged figure, $3,850 in both 2025 and 2026. All three memory platforms gave $3,500 instead, in ten responses each, thirty in total. That was the 2024 figure. It changed between 2024 and 2025 and we had recorded it only as unchanged for the year in question. It turned out to measure something more useful: how far back a platform's figure comes from. We report it as an instrument note rather than removing it.
A prior-year figure that was itself wrong. Asked about the standard deduction for married couples filing jointly, Claude gave 2025 as $30,000 in ten responses and Gemini in three. The IRS puts tax year 2025 under the One Big Beautiful Bill Act at $31,500. A statute changed the number mid-year, and the platforms' confidently stated last-known figure was superseded before it was ever current. Perplexity and Google AI Mode gave the correct 2026 figure of $32,200 in ten responses each.
What this does not show
- Fourteen figures on one date. Selected, not a census. No inference to federal thresholds generally.
- Counts, not rates. Five asks bounds run-to-run variance; it does not establish a rate.
- We measured provider APIs, deliberately. API surfaces are controlled, repeatable and auditable. A frozen corpus can be re-read by anyone; a scraped session cannot. Consumer applications carry personalization, location and session state that API calls do not, and published guidance in this field indicates they may behave differently. That difference is a property of this method, stated here rather than left implicit.
- We did not measure whether anyone was harmed. We measured what was said.
- Classification is automated and then read. Every threshold was reviewed by hand, which is how all three corrections above were found.
- One platform required a higher token ceiling than the others. All affected responses were re-collected rather than excluded; see the corrections section.
What publishes next
In September, two studies:
- Ten Triangle accounting firms performing federal Single Audit work, drawn from the Federal Audit Clearinghouse: every firm named as auditor on three or more Triangle engagements. A complete population, not a sample. It records what five AI platforms said when asked the questions a nonprofit board asks, checked against the filings that name who performed each audit. The threshold finding above is one paragraph of it.
- 15 registered investment advisers, drawn at random from the SEC's register, with the questions derived from each firm's own Form ADV. Not what AI says when you type a firm's name, but what it says when a buyer describes exactly the service that firm provides, without naming it.
Those are different questions, and the gap between them is the finding.
Both studies publish with the roster, the questions as sent, the frozen artifacts and their fingerprints, and a record of every correction made. Every measured firm receives its own complete record before publication.
Corpus
| Artifact | Detail |
|---|---|
| Instrument | 14 thresholds, 2 prompts, frozen and hashed before collection |
| Ground truth | Verified against IRS releases, Revenue Procedures and the CFR, August 20, 2026 |
| Corpus | 700 events, 0 collection errors, 0 empty responses, 0 truncated, frozen and hashed |
| Re-collection | 4 Gemini events returned HTTP 503; 93 Gemini events were truncated by the token ceiling. All re-collected the same day with prompt text, platform and ask index unchanged. Superseded corpora retained. |
| Scoring | Published scoring code applied to the frozen corpus; output frozen and hashed, carrying the instrument and corpus fingerprints |
| Platforms | ChatGPT · Claude · Perplexity · Gemini · Google AI Mode |
| Collection date | August 20, 2026 |
| Models measured | gpt-4o-2024-08-06 · claude-opus-4-5-20251101 · gemini-3.6-flash · perplexity sonar · Google AI Mode via SerpAPI. Resolved strings recorded per event. We do not assert training cutoffs for these models. |
The questions as sent, the artifacts and their fingerprints, the collection conditions and the ground truth are published in full.
Corrections and disputes
If we have described a figure or a platform's behavior incorrectly, tell us. We will check it against the source, publish a dated correction if we were wrong, and say so plainly. We correct errors of fact. We do not change findings because someone dislikes them.
AIWorthy is an independent research firm measuring how AI systems describe companies against the public record. We publish what we find. We sell measurement, and we sell no optimization, content, or remediation of any kind.
hharreld@aiworthy.ai · aiworthy.ai