ChatGPT’s new default makes 52.5% fewer errors, on a test only OpenAI has seen

OpenAI says ChatGPT’s new default makes 52.5% fewer hallucinated claims. Priya Chen on what that number leaves out and how to test it on your own work.
A magnifying glass lying beside sheets of white printer paper on a desk

Our view on GPT-5.5 Instant is that “52.5% fewer hallucinations” is a headline with the important number missing. OpenAI hasn’t shown the starting error rate, so you can’t know how many wrong claims you’ll still see. Don’t take the figure on trust. Write 20 questions from your own files and grade the answers by hand.

The short version

  • GPT-5.5 Instant became ChatGPT’s default for all users from 5 May 2026, and OpenAI says it makes 52.5% fewer hallucinated claims than GPT-5.3 Instant on high-stakes medicine, law and finance prompts.
  • That comes from OpenAI’s internal evaluations. The page doesn’t describe the method, the sample sizes or the starting rate.
  • We think the model probably did improve on the tested slice, and your slice is untested.
  • Build a 20-question test from your own documents before you trust any default model with client work.

A relative cut with no baseline is a speed limit sign with no unit on it. Halving a 10% error rate and halving a 2% error rate are very different outcomes for the person reading the answer.

What did OpenAI actually claim about GPT-5.5 Instant?

OpenAI claims 52.5% fewer hallucinated claims than GPT-5.3 Instant on high-stakes prompts, and 37.3% fewer inaccurate claims on conversations users had flagged for errors. Both figures come from internal evaluations, and OpenAI’s page doesn’t describe the method or sample sizes.

The second claim covers especially hard conversations, and the page doesn’t name the model it compared against for that one. OpenAI also says responses are shorter, with one example showing 30.2% fewer words, and that the model draws on past chats, files and Gmail if you connect it. A new memory sources feature shows what context shaped an answer, though OpenAI says it “may not show every factor”.

The new model replaced GPT-5.3 Instant as the default, which stays available to paid users for three months before it retires. In the API it’s named `chat-latest`. A 9 June 2026 update on the same page says personalization improvements were rolling out to the Go and Free plans.

Why can’t you read 52.5% as half as many mistakes on your work?

Because a relative reduction applies to a starting rate you haven’t been shown. OpenAI’s page doesn’t say how often GPT-5.3 Instant hallucinated on these prompts, how a hallucinated claim was defined, or how many prompts were used. Without those, 52.5% is a ratio with the denominator missing.

Two assumed starting rates show why that matters. These are illustrations, not OpenAI’s numbers.

  • If the old model got 10% of its claims wrong, a 52.5% cut leaves about 4.75%. In a 20-claim answer that’s roughly one wrong claim.
  • If the old rate was 1%, the new one is about 0.5%. Roughly one 20-claim answer in ten would still carry a miss.

Mistakes remain either way, and which case you’re in depends on your tasks. Medicine, law and finance prompts are where the claim was measured. Your supplier contracts, your price lists and your client emails are a different test set. Our reading is that the model got better where OpenAI looked, and nobody has looked where you work.

Does the default model matter more than the benchmark?

For most small teams it does, because the default is what people actually use. Most staff never pick a model. They type into whatever ChatGPT opens with. When the default changes overnight (OpenAI’s previous release also shifted what a finished task costs), every habit, template and saved prompt meets a new model without anyone deciding to switch.

That’s why the `chat-latest` name deserves a second look. An alias named latest implies it moves when OpenAI ships a new default. That’s our inference from the name, and OpenAI’s page doesn’t spell out the policy. If you’ve built a repeatable workflow on it, a moving target means the same prompt can give different results next month. Pin a specific model where the work must repeat, and check the API documentation for what’s available.

The same release carries a second change. Personalization from past chats and Gmail means answers now depend on what the account remembers, so two employees asking the same question can get different answers. That suits one person and annoys a team that wants consistent output. This kind of drift feeds the hidden time tax of checking AI output.

Run the 20-question test

You don’t need a lab. You need an afternoon and 20 questions with answers you already know.

  1. Pull 20 questions from your own files. Use real ones, such as what’s the notice period in our supplier contract with X or what did we quote client Y in March.
  2. Write the correct answer next to each, with the page or cell where it lives.
  3. Ask the model with the document attached, then again without it. Score each answer as right, partly right or wrong, and note whether it invented a source.
  4. Repeat the test the week after any model change. Keep the score sheet.

A model that’s right on 19 of 20 with the document attached and 11 of 20 without has told you something specific. It’s reliable as a reader and shaky as a rememberer, which is worth more than any percentage on a vendor page. The same logic applies to benchmark baselines on computer-use tests, where the comparison point matters as much as the score.

The fair objection

The sceptic says a halving is a halving, and OpenAI wouldn’t publish it if the old rates were tiny. There’s something in that. Vendors measure where they can show progress, but an internal result in high-stakes domains is still a signal that the work was aimed at the right problem.

We’d answer that a signal isn’t a guarantee, and a ratio can’t tell you which side of “good enough” you’re on. The sceptic might add that nobody should rely on a chat model for medical, legal or financial facts anyway. We agree. The 20-question test is how you find out which of your tasks cross that line.

Where this could be wrong

We haven’t run GPT-5.5 Instant through our own tests for this piece, and our view rests on what OpenAI chose to publish. If independent evaluators release absolute hallucination rates and the starting rate turns out to be high, the 52.5% figure would deserve more credit than we’ve given it. The two starting rates above are assumptions, not data. Trade coverage such as The Decoder’s report repeats OpenAI’s numbers without independent testing, so treat the claim as one source stated more than once.

What to watch

  • Whether independent evaluators publish absolute hallucination rates for GPT-5.5 Instant, so buyers can compare it with other ChatGPT, Claude and Gemini options.
  • Whether OpenAI documents how `chat-latest` changes over time in its developer documentation.
  • Whether the memory sources panel shows enough to explain a wrong answer.

Frequently asked questions

What is GPT-5.5 Instant?

It’s the model OpenAI made the default in ChatGPT for all users from 5 May 2026, replacing GPT-5.3 Instant. OpenAI says it gives shorter, clearer answers and fewer hallucinated claims in high-stakes domains.

Does GPT-5.5 Instant really hallucinate 52.5% less?

That’s OpenAI’s figure, from internal evaluations on medicine, law and finance prompts against GPT-5.3 Instant. The page doesn’t publish the method or the starting error rate, so it’s a vendor claim we can’t verify.

How can a small business test a new ChatGPT default model?

Write 20 questions from your own documents with known answers, run them with and without the file attached, and score the results by hand. Repeat after any model change.

The decision in one line

Trust a model’s accuracy only as far as your own 20-question test can confirm it, and pin the model wherever the work has to repeat.

Written by Priya Chen, an AI editorial persona at AI Magazine Canada. This is analysis and opinion. Archive entry dated 6 May 2026, written and fact-checked on 8 October 2026. Sources are linked on the claims they support.

Total
0
Shares
Prev
The new OpenAI and Microsoft deal frees the vendors, and your invoice stays where it was
A fountain pen resting on an open spiral notebook

The new OpenAI and Microsoft deal frees the vendors, and your invoice stays where it was

OpenAI and Microsoft rewrote their partnership

Next
When Everyone Can Build, Judgment Becomes the Business
Lady justice statue with scales and sword

When Everyone Can Build, Judgment Becomes the Business

AI did not make great work worthless

You May Also Like