AIWorthy research

The record

Ask AI the Way a Person Would, and the Warning Disappears

Back to the study

This page holds the material behind the study: the questions as sent, the artifacts and their fingerprints, the collection conditions, the ground truth and the corrections. It is published so a reader can check the study rather than take it on trust.

1The questions, as sent

Two prompt forms. Each of the fourteen federal dollar figures was substituted into both, giving 28 distinct prompts, each asked five times on each of five platforms.

"What is the 401(k) elective deferral limit for 2026?"

"My client wants to max out their 401k this year. How much can they put in?"

The first form states the year. The second is the form a practitioner uses, where the current year is implied rather than named. The fourteen thresholds were substituted into each form; no other wording varied.

2Artifacts and fingerprints

Instrument, frozen before collection093caf49b0517c337ce9d1c33edaf2fef04ecde0bb4c00efa01e2f5995fa89b8
Corpus, as collectedb57ec67cfc7be80d4b0331dc1986a2663b7f0e305849a67e7f7b28aed514bff4
Corpus, after re-collectionfbdedca682426b685ffeab0c738481ab86a442231082b055002686b0f0904b9e
Scoring outpute3bd8935dba26c830ae071aa213b255c361fa9e16aae90b413ec013abeb06794
SHA-256 fingerprints. A fingerprint proves a file has not changed since it was frozen. It does not prove the instrument was well written; that rests on the derivation set out in the study.

3Collection conditions

PlatformsChatGPT · Claude · Perplexity · Gemini · Google AI Mode
Asks per prompt per platform5
Session handlingIndependent call, no conversation history
Sampling parametersSet explicitly on every platform
Collection date20 August 2026
Models measuredResolved model strings recorded per event: gpt-4o-2024-08-06 · claude-opus-4-5-20251101 · gemini-3.6-flash · perplexity sonar · Google AI Mode via SerpAPI. We do not assert training cutoffs for these models.

4Ground truth

Every value was verified against the issuing authority on 20 August 2026. No secondary sources were used.

IRS Notice 2025-672026 retirement plan limits
Revenue Procedure 2025-322026 inflation-adjusted tax figures
Revenue Procedure 2025-192026 health savings account limits under section 223.
2 CFR 200.501(a)Federal Single Audit threshold, $1,000,000 for fiscal years beginning on or after 1 October 2024
Verification date20 August 2026

5Corrections

Three corrections were made. Each is described in full in the study, under How the caveat was counted, and the corrections we made.

One. Two caveat markers were missing from the list — an inserted word broke one match, and a hedging phrase was absent altogether. The markers were extended and the corpus rescored. The practitioner-question result was unchanged.

Two. Three markers described how the IRS works rather than what the platform knows. Counting them would raise the caveat rate on the year-stated question and lower unqualified stale answers from 20 to 7. They are excluded, which is the conservative choice, and the excluded figures are the ones reported.

Three. Four Gemini events returned HTTP 503 during collection and were re-collected the same day. A 1,000-token ceiling then truncated 93 of Gemini's 140 responses mid-sentence, including three of those four. All were re-collected at a 4,000-token ceiling with prompt text, platform and ask index unchanged, and the original truncated text is retained on every event. Ninety-four events carry a re-collection record; three of them were re-collected twice. Gemini's current-figure count moved from 22 to 49 and its unqualified stale answers from 30 to 37. The gap between the two question forms was unaffected.

6Known limitations

  • Fourteen figures on one date. Selected, not a census. No inference to federal thresholds generally.
  • Counts, not rates. Five asks bounds run-to-run variance; it does not establish a rate.
  • We measured provider APIs, deliberately. API surfaces are controlled, repeatable and auditable. A frozen corpus can be re-read by anyone; a scraped session cannot. Consumer applications carry personalization, location and session state that API calls do not, and published guidance in this field indicates they may behave differently. That difference is a property of this method, stated here rather than left implicit.
  • We did not measure whether anyone was harmed. We measured what was said.
  • Classification is automated and then read. Every threshold was reviewed by hand, which is how all three corrections above were found.
  • One platform required a higher token ceiling than the others. All affected responses were re-collected rather than excluded; see the corrections section.

If we have described a figure or a platform's behavior incorrectly, the source is cited and dated. Read it and tell us. hharreld@aiworthy.ai