OpenAI announced GPT-5.5 on April 23, 2026, positioning it as the successor to GPT-5.4 and the company's new flagship model across ChatGPT and the API. On May 5, 2026, a version called GPT-5.5 Instant began rolling out as the new default model for all ChatGPT users, including the free tier, replacing GPT-5.3 Instant as the model answering everyday queries. A further quality update to GPT-5.5 Instant followed on May 28, 2026, tuning response style rather than changing the underlying architecture.
Instant and Thinking, not one model
GPT-5.5 is not a single model but a routed system: prompts that need speed go to the Instant path, while prompts that benefit from multi-step reasoning are routed to Thinking, with ChatGPT choosing automatically unless a user overrides it. Both variants are still built on the transformer architecture that underlies every GPT model, but OpenAI has trained the router itself — deciding when to think longer, when to call a tool, and how to structure multi-step answers — as a learned behavior, which is what differentiates GPT-5.5 from a simple bigger version of GPT-5.4.
Long context: the less flashy but more practical improvement
Beyond the headline agentic benchmarks, OpenAI's own release data shows GPT-5.5 making large gains on long-context retrieval tasks, which matter more for everyday use than they might sound. On OpenAI's internal MRCR v2 test — which checks whether a model can accurately retrieve a specific "needle" of information buried in a huge amount of surrounding text — GPT-5.5 scores 74.0% at the 512K-to-1-million-token range, compared with 36.6% for GPT-5.4. That's the kind of improvement that shows up when someone pastes a very long document, codebase, or chat history into the model and asks a specific question about something buried deep inside it, rather than in short back-and-forth chat.
Where the benchmark numbers actually come from
OpenAI's own published comparison table for GPT-5.5, dated the day of the April 23 announcement, credits the model with 82.7% on Terminal-Bench 2.0 (a test of agentic command-line work), versus 75.1% for GPT-5.4, 69.4% for Anthropic's Claude Opus 4.7, and 68.5% for Google's Gemini 3.1 Pro. On GDPval — OpenAI's benchmark for economically valuable knowledge work across 44 occupations — GPT-5.5 scores 84.9%, ahead of Opus 4.7's 80.3% and Gemini 3.1 Pro's 67.3%. On OSWorld-Verified, a test of controlling a real desktop environment, GPT-5.5 reaches 78.7%, edging out Opus 4.7's 78.0%. These figures come directly from OpenAI's own release materials and have been reproduced in independent write-ups from outlets including CNBC, TechCrunch, and ZDNet, though it's worth flagging that OpenAI is grading its own model on benchmarks it partly helped popularize, so treat the exact percentages as OpenAI-reported rather than independently audited.
Not every number favors GPT-5.5. On SWE-Bench Pro, a public coding benchmark, GPT-5.5 scored 58.6%, behind Claude Opus 4.7's 64.3%. Independent analysis site Artificial Analysis has also flagged that GPT-5.5 posts one of the highest hallucination rates recorded on its AA-Omniscience benchmark alongside its highest-ever accuracy score, a reminder that raw capability gains and reliability don't always move together.
What actually changed for users
For most ChatGPT users, the practical change is that GPT-5.5 Instant now handles the majority of everyday conversations by default, with a noticeably different response style that OpenAI adjusted again in its late-May update after user feedback. For developers and businesses, GPT-5.5 is pitched primarily at agentic coding and long-horizon computer-use tasks — writing and running code across multiple steps, operating tools and terminals, and handling long documents, where OpenAI's MRCR long-context benchmarks show meaningful gains over GPT-5.4 at context windows beyond 256,000 tokens.
- Announced April 23, 2026; GPT-5.5 Instant became ChatGPT's default on May 5, 2026
- 82.7% on Terminal-Bench 2.0, ahead of GPT-5.4 (75.1%) and Claude Opus 4.7 (69.4%)
- 84.9% on GDPval knowledge-work benchmark; 78.7% on OSWorld-Verified computer-use benchmark
- Trails Claude Opus 4.7 on SWE-Bench Pro (58.6% vs 64.3%)
Our take
GPT-5.5's headline wins are real and well-documented, but they're concentrated in agentic and computer-use tasks that OpenAI itself chose to highlight — this is a strong model for coding agents and long document work, not a clean sweep of every benchmark that exists. The gap between GPT-5.5's high accuracy and its elevated hallucination rate on independent testing is the detail worth watching if you're deploying it for anything where factual precision matters more than task completion.
What to watch next
Watch for OpenAI's next numbered release (already previewed under the name GPT-5.6 in some of OpenAI's own materials) and for independent, non-OpenAI-run benchmark results that test GPT-5.5 against Gemini and Claude's newest models under neutral conditions rather than OpenAI's own comparison tables. It's also worth watching whether OpenAI or independent researchers publish more detail on the hallucination-rate tradeoff that Artificial Analysis flagged, since that has direct implications for anyone deploying GPT-5.5 in customer-facing or compliance-sensitive settings rather than agentic coding tasks where a wrong answer is quickly caught and corrected.