First look · Model choice

Claude Fable 5 vs GPT-5.5: the benchmark lead comes with a safeguard question

Claude Fable 5 looks like a serious step up for long-horizon coding and agent work. But the buyer question is not just “which model is smarter?” It is whether your real task gets the full Fable 5 behavior, what it costs, and how often safeguards change the experience.

My bottom line first: Fable 5 may be the more capable long-horizon coding and agentic model on paper. GPT-5.5 may still be the safer default for teams that care about predictable behavior, lower API cost, and fewer surprises from domain safeguards.

Why this comparison matters today

Anthropic just put Claude Fable 5 in public hands, and the first reaction is predictable: people want to know whether it beats GPT-5.5. That is a reasonable question, but it is too flat.

The more useful version is: if Fable 5 looks stronger on benchmarks, will your actual workflow reliably get that strength? Anthropic says Fable 5 is its most capable widely released model and that the longer the task, the larger its lead over previous Claude models. It also says the public model uses new safeguards, and some cybersecurity, biology, chemistry, or distillation-related requests are handled by Claude Opus 4.8 instead. See AIAnthropic launch note and AIClaude model overview.

That makes this a buyer note, not just a benchmark note.

What Claude Fable 5 actually is

Anthropic describes Claude Fable 5 as a Mythos-class model made safe enough for general release. Claude Mythos 5 uses the same broad capability class but is limited to Project Glasswing-style trusted access. Fable 5 is the public version with additional classifiers and fallback behavior.

On the API side, Anthropic lists claude-fable-5 with a 1M token context window, 128k max output, always-on adaptive thinking, and pricing at $10 per million input tokens and $50 per million output tokens. Anthropic’s pricing page also shows batch pricing at $5 input and $25 output per million tokens. AIClaude model overviewAIClaude API pricing

GPT-5.5 is cheaper on the standard API: OpenAI’s model page lists a 1,050,000 token context window, 128k max output, and $5 input / $30 output per million tokens. OpenAI’s launch post also says GPT-5.5 is available in ChatGPT and Codex, with API pricing at the same $5 / $30 rate. OGPT-5.5 API model pageOOpenAI GPT-5.5 launch post

The benchmark story

I would read the first-day benchmark story in two layers. First, the direction is clear: Fable 5 is being positioned for hard software engineering, knowledge work, vision, memory, and long-running agentic tasks. Second, some exact benchmark details are still harder to compare cleanly because they come from provider-run tables, embedded charts, partner quotes, or different harnesses.

Key benchmark scores: Fable 5 vs GPT-5.5

The table below is a first-look score summary, not a final verdict. Some Fable numbers come from Anthropic launch material or third-party transcriptions of Anthropic charts, and starred rows should be read with the public-Fable safeguard caveat in mind. AIAnthropic launch noteKKingy AI benchmark transcription

BenchmarkClaude Fable/Mythos 5GPT-5.5My read
SWE-Bench Pro80.3%58.6%Fable 5 looks much stronger on complex software engineering tasks.
FrontierCode Diamond29.3%5.7%This is the biggest-looking gap and supports the long-horizon coding story.
OSWorld-Verified85.0%78.7%Fable leads, but GPT-5.5 still looks strong.
GDP.pdf vision29.8%24.9%Fable appears somewhat better on document and vision understanding.
Humanity’s Last Exam, with tools64.5%*52.2%Strong result, but the starred caveat matters.
Terminal-Bench 2.188.0%*83.4%Fable/Mythos appears ahead, but harness differences matter.
AutomationBench17.4%12.9%Both are still early for automation agents.
Legal Agent Benchmark13.3%2.1%Fable looks ahead, but legal-agent capability is still limited overall.
HealthBench Professional66.0%*51.8%Potentially strong, but public Fable may be more constrained in health or bio-adjacent use.
AreaWhat the sources sayHow I would read it
Software engineeringAnthropic says Fable 5 scores highest on Cognition’s FrontierCode eval and reports strong early results from partners such as Stripe. OpenAI reports GPT-5.5 at 58.6% on SWE-Bench Pro and 82.7% on Terminal-Bench 2.0.Fable 5 looks strongest when the task is long, messy, and repo-level. GPT-5.5 is still strong, but this is the area where Fable 5’s launch case is most convincing.
Knowledge workAnthropic says Fable 5 is the first to break 90% on one partner’s core analytics benchmark. OpenAI reports GPT-5.5 at 84.9% on GDPval and 60.0% on FinanceAgent v1.1.Both models are credible here. I would test on your own documents before switching, because finance/legal/analytics evals are sensitive to prompt format and source quality.
Vision and computer useAnthropic calls Fable 5 state-of-the-art for vision tasks and gives examples such as game playing from raw screenshots. OpenAI reports GPT-5.5 at 78.7% on OSWorld-Verified and 83.2% on MMMU Pro with tools.Fable 5 may be better at long visual-agent tasks. GPT-5.5’s public numbers still make it a serious baseline for computer-use workflows.
Context and outputAnthropic and OpenAI both list roughly 1M-token API context and 128k max output for these API models.The spec line is similar. The buying question moves to quality, cost, tool behavior, and how often long context actually helps.
CostFable 5 API list price is $10 input / $50 output per million tokens. GPT-5.5 standard API list price is $5 input / $30 output per million tokens.Fable 5 has to win by enough on task completion or token efficiency to justify the higher unit price.

The headline is not “Fable destroys GPT-5.5.” The more careful headline is: Fable 5 looks ahead in the places where long-horizon agent work matters most, but the cost and safeguard layers are part of the real comparison.

The safeguard story

This is the part I would not bury. Anthropic says Fable 5 uses classifiers for cybersecurity, biology and chemistry, and distillation-related requests. If a classifier triggers, the response is automatically handled by Claude Opus 4.8, and the user is informed. Anthropic says more than 95% of Fable sessions involve no fallback, and those sessions should perform effectively like Mythos 5. AIAnthropic launch note

That is a reasonable safety design. It is also a product behavior that buyers need to test. A security team, biology researcher, malware analyst, LLM researcher, or even a developer working on security-adjacent code may care less about the average fallback rate and more about fallback rate inside their own task category.

Business Insider also reported that some safeguards may block or reroute benign requests in sensitive areas, including examples around biology/cancer questions. That is a useful reminder: false positives are not an edge case if your work lives near the boundary. BIBusiness Insider safeguard report

What people are saying so far

The first-day community signal is split. Some early testers describe Fable 5 as a real step change, especially for difficult agentic work. A Hacker News impressions thread says the model felt meaningfully stronger, while also noting that classifiers were very sensitive on some benign coding tasks. YHN impressions thread

Other feedback is more skeptical. In the main Reddit launch thread, users focused on pricing/access after June 22 and aggressive safety filtering. One user reported being switched off Fable while asking for help with a catered-meal shopping list; another said Opus 4.8 already felt good enough for their coding workflow. rr/ClaudeAI launch thread

There is also a separate Reddit benchmark discussion arguing that the picture is more mixed outside Anthropic’s strongest categories, including finance and Vending-Bench research notes. I would treat that as market signal rather than settled fact, but it is exactly the kind of counterweight a buyer should read before switching a workflow. rr/ClaudeAI benchmark thread

On Hacker News, one recurring worry is transparency: what happens when safeguards change model behavior, and can users tell? That discussion is less about raw intelligence and more about whether a high-capability model remains predictable inside real work. YHN safeguard transparency thread

Where Fable 5 looks better than GPT-5.5

I would test Fable 5 first when the work is long-horizon and hard to decompose:

  • large repo migrations or multi-file refactors;
  • agentic coding tasks where the model has to plan, inspect, run, and revise;
  • visual reasoning that spans screenshots, documents, or UI reconstruction;
  • deep analysis where the model has to maintain a thread across many turns or many files;
  • high-value tasks where a higher unit token price is acceptable if fewer attempts are needed.

The key phrase is “high-value.” At $10 / $50 per million tokens, Fable 5 should not be the default for every chat or every code question. It should earn its place on the hardest tasks.

Where GPT-5.5 still looks safer to choose

I would keep GPT-5.5 as the default when predictability and budget matter more than the top-end ceiling:

  • daily research, writing, and analysis work;
  • workflows already tuned for ChatGPT, Codex, or the OpenAI API;
  • cost-sensitive production systems where output tokens dominate the bill;
  • tasks near safety-boundary topics where Fable fallback could interrupt the run;
  • teams that would rather have a slightly lower ceiling than a less predictable routing layer.

OpenAI’s ChatGPT help article also gives clearer tier behavior for GPT-5.5 Instant, Thinking, and Pro, including context-window differences by tier. OOpenAI Help: GPT-5.5 in ChatGPT

My switching test

I would not switch because of one benchmark table. I would run a small test:

  1. Pick three tasks that failed or took too long on GPT-5.5.
  2. Run them on Fable 5 with the same instructions and source material.
  3. Record total tokens, wall-clock time, number of retries, and whether fallback happened.
  4. Check the output with the same human review standard.
  5. Switch only if the final cost per accepted result improves.

The last line matters. A more expensive model can be cheaper if it finishes the work with fewer loops. It can also be more expensive in exactly the boring way: every attempt costs more, and the win rate does not improve enough.

My bottom line

Claude Fable 5 may be the most interesting public coding and agentic model release right now. The early evidence points toward a higher ceiling on long, complex work. But the public product is not just “Mythos capability.” It is Fable 5 plus safeguards, fallback behavior, access rules, and a higher API price.

That does not make GPT-5.5 the winner. It makes GPT-5.5 the more boring default. And for many teams, boring is valuable.

My first-look recommendation: test Fable 5 on the hardest workflows where GPT-5.5 is already costing you time. Keep GPT-5.5 as the default until Fable 5 proves that its higher ceiling survives your task, your budget, and your safety boundary.

More Research Notes