AI / Comparison
Opus 5.5 against GPT-6 on the work developers actually do
The verdict
For the agent you leave running - a refactor across forty files, a two-hour session, a migration you review rather than write - take Claude Opus 5.5. Anthropic positions it explicitly for long-running agentic coding, it undercuts the GPT-6 frontier tier badly on output tokens at 20 dollars against 50, and its reliable knowledge cutoff of June 2026 matters more than a benchmark point when you are working with a library that shipped six months ago. For breadth across a platform, and for anything cost-sensitive at volume, take GPT-6. The 6.1 Sol tier is competitive on agentic coding and now undercuts Opus 5.5 itself at 2 dollars in and 10 out, and GPT-6 Luna goes lower still, which decides any high-throughput workload on price alone.
Every figure on this page was checked against Anthropic's and OpenAI's own documentation in October 2026. Both families moved again in the last six weeks - Claude Opus 5.5 shipped September 22 and Sonnet 5.5 September 28, 2026; GPT-6 Sol and Luna arrived September 22 and GPT-6.1 Sol on September 29 - so assume anything undated elsewhere is stale.
Claude vs GPT: the published specifications
Prices are per million tokens on the standard synchronous API, as listed by each vendor in September 2026. Batch and caching discounts apply on both sides and are not shown.
| Dimension | Claude Opus 5.5 | GPT-6.1 Sol |
|---|---|---|
| Released | September 22, 2026 | September 29, 2026 (GPT-6 Sol and Luna arrived September 22) |
| Positioned for | Long-running agentic coding and knowledge work | Complex professional work; OpenAI's stated best coding model |
| Price per MTok | 4 dollars in, 20 dollars out | 2 dollars in, 10 dollars out |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128k tokens, or up to 300k on the Batch API with a beta header | 128k tokens |
| Reliable knowledge cutoff | June 2026 | April 2026 (May 2026 for GPT-6 Luna) |
| Reasoning control | Adaptive thinking, always on and no longer disableable, plus an effort parameter that defaults to medium on this model | Reasoning effort selectable per request and from the Codex CLI |
| Cheaper siblings | Sonnet 5.5 at 2 dollars in and 10 out on a 1M context; Haiku 4.5 at 1 and 5 | GPT-6 Luna at 10 cents in and 50 cents out, on the same 1.05M context |
| Halo model above it | Claude Fable 5.1, generally available September 1, 2026, at 10 dollars in and 50 out | GPT-6 Astra, generally available in early September 2026, at 10 dollars in and 50 out |
Gotcha: the context windows look identical and are not comparable token for token. Anthropic notes that Opus models from 4.7 onward, Opus 5.5 included, use a new tokenizer that may use roughly 1x to 1.35x as many tokens for the same text - up to about 35 percent more, varying by content. Measure your own corpus rather than assuming a 1M window holds the same repository on both sides.
What the benchmarks claim, and how much to believe
Treat every number below as a claim with a date, not a property. Agentic coding scores depend on the harness as much as the model, and the published leaderboards moved several times during 2026.
Anthropic's own Opus 5.5 claims, September 22, 2026. This launch post does lead with head-to-head numbers, and it aims them at OpenAI's top tier: 66.4 percent on Terminal-Bench 4.0 at xhigh effort against 57.9 percent for GPT-6 Astra at high effort, and 54.4 percent on FrontierCode v1.1 against Astra's 53.3. Against its own family, Anthropic reports 52.5 percent on CursorBench 4.0 at default effort, ahead of Fable 5.1 (51.8) and Opus 5 (46.6) at max effort, and says Opus 5.5 costs 40 percent less than Opus 5 on typical workloads while generating output more than 30 percent faster. The effort setting attached to each figure is part of the claim.
Anthropic's own Opus 5 claims, July 24, 2026. The launch post is notable for what it does not lead with. There is no SWE-bench headline. Instead Anthropic quotes agentic and computer-use evaluations: Frontier-Bench v0.1, where Opus 5 more than doubles Opus 4.8's performance at a lower cost per task; CursorBench 3.2, where it lands within 0.5 percent of Fable 5's peak score at half the cost per task; ARC-AGI 3, where it scores three times the next-best model; OSWorld 2.0, where it surpasses Fable 5's best result at just over a third of the cost; and Zapier's AutomationBench, where the pass rate is around 1.5 times the next-best model. Cost per task appears in almost every one of those sentences, which is the story.
OpenAI's GPT-5.6 claims, July 9, 2026. Sol is described as the strongest coding model in the family and as OpenAI's strongest cybersecurity model, with independent aggregators reporting Terminal-Bench 2.1 results in the high eighties for Sol and higher still for its maximum-effort configuration. Third-party coding-agent indices published this summer put GPT-5.6 Sol running inside Codex at or near the top of their composite rankings.
The figures circulating that neither vendor published. You will see "96 percent on SWE-bench Verified" attributed to Opus 5 across a lot of summary sites. That number comes from third-party aggregation, not from Anthropic's announcement. It may well be accurate; it is not a vendor claim, and it should not be quoted as one. The same caution applies to any single Terminal-Bench figure - which variant, which effort level, and which harness all move it by several points.
The honest summary: on frontier agentic coding these two are close enough that the harness, the prompt, and your repository's test suite matter more than the model choice. Where they are genuinely not close is the price ladder below the flagship, and that is where most real spending happens.
When each family is the right call
Choose Claude when
- The workload is a long-running coding agent. This is the use case Anthropic optimizes for and markets against, and Claude Code's adoption among senior developers reflects it.
- You are working with libraries and APIs that changed in the first half of 2026. Opus 5.5's reliable knowledge cutoff is June 2026.
- Output tokens dominate your bill and you want frontier quality. At 20 dollars against 50 per million out, sustained code generation is far cheaper on Opus 5.5 than on GPT-6 Astra - though GPT-6.1 Sol undercuts them both at 10.
- You need very large single responses. The Batch API supports up to 300k output tokens behind a beta header, which no comparable tier here matches.
- You want a step above the workhorse without leaving the family. Fable 5.1 exists at 10 and 50 per million, with cache reads cut to 25 cents on September 1, 2026 - which matters more than the sticker price on agentic work that re-reads the same context.
Choose GPT when
- Volume is the constraint. Luna at 10 cents in and 50 cents out on a 1.05M context has no equivalent on the Anthropic price list, and it decides any high-throughput classification, extraction, or summarization workload outright.
- You want one price ladder with one context window. Astra, Sol, and Luna all take 1.05M tokens, so you can downgrade a call to a cheaper tier without redesigning the prompt.
- Your team already pays for ChatGPT. Codex is bundled with the subscription rather than sold separately, which removes a procurement conversation entirely.
- The task needs OpenAI's wider tool surface - hosted web search, file search, and computer use available through one API alongside the model.
- You need the cheapest capable mid-tier. GPT-6.1 Sol and Sonnet 5.5 are both listed at 2 dollars in and 10 out, so price no longer separates them and the decision moves back to behavior.
- You want OpenAI's own top tier rather than its workhorse. GPT-6 Astra reached general availability in early September 2026 at 10 and 50 per million on a 1.05M context, which prices it level with Fable 5.1 and at two and a half times Opus 5.5 on output. Anthropic's launch numbers put Opus 5.5 ahead of Astra, but those are vendor claims at chosen effort settings, so treat it as a close race rather than a settled one.
In practice: agent behavior beats benchmark scores
The thing that decides whether a two-hour agent session succeeds is not raw capability. It is whether the model asks before doing something irreversible, whether it notices its own failed command instead of confidently proceeding, whether it stays coherent after compaction, and whether it stops when it is stuck rather than trying the same fix a fourth time. None of those are on a leaderboard, and all of them differ between families and between harnesses.
The practical consequence is that you should test on your own repository before committing. Give both the same task with the same rules file - a real one, from your backlog, that touches at least four files - and score them on the review, not the demo. How much did you have to correct? What did it delete? Did it invent a dependency? That afternoon is worth more than every benchmark on this page.
Also worth saying plainly: you do not have to pick one forever. Cursor exposes models from multiple vendors, Copilot gates premium models by tier, and open-source agents like Aider and Cline point at any endpoint you give them. Claude Code is the exception - it runs Anthropic models only - which is a genuine consideration if multi-vendor flexibility is a requirement rather than a preference.
Where this fits
Models are the layer underneath the tools. For the tool decision, read Claude Code vs Codex, which pairs each model family with its native terminal agent, or the full assistant field guide for all ten options. The workflow guide covers matching model tier to task, which is the single biggest lever on cost.
Building AI features into a product rather than using AI to write it? The AI APIs and SDKs directory covers the client libraries, gateways, and vector stores, and the AI app stack guide assembles a full recommendation. Official model documentation lives at platform.claude.com and developers.openai.com.