Claude Opus 5.5 vs. GPT-6: Benchmarks, Falling Prices, and the Push for a Shared AI Safety Standard
A fact-checked comparison of Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Astra: benchmarks, pricing, safety restrictions, the coming Anthropic and OpenAI IPOs, and the plan for a shared AI safety standards body.

September 2026 may be remembered as the month the frontier got crowded, and cheaper. OpenAI released GPT-6 Astra on September 3. On September 22 it followed with the smaller GPT-6 Sol and GPT-6 Luna, and Anthropic released Claude Opus 5.5 the same day.
The launch timing is only part of the story. Both companies are preparing to go public, both are cutting prices on the way there, and this week they were reported to be moving ahead, together with Google, on a shared AI safety standards body. That is a lot for one month, so we will take it in order.
A note on naming: There is no product called "ChatGPT 6.0." ChatGPT is OpenAI's app, and GPT-6 is the model family behind it. GPT-6 comes in three sizes: Astra (the flagship), Sol (mid-tier), and Luna (small and inexpensive). This article compares Opus 5.5 mainly with Astra, and brings in Sol where price is the question.
At a Glance
| Claude Opus 5.5 | GPT-6 Astra | GPT-6 Sol | |
|---|---|---|---|
| Released | Sept 22, 2026 | Sept 3, 2026 | Sept 22, 2026 |
| API price per 1M tokens (input / output) | $4 / $20 | $10 / $50 | $2 / $10 |
| Context window | 1M tokens | ~1.05M tokens | ~1.1M tokens |
| Max output | 128K tokens | 128K tokens | Not stated |
| Where to get it | Claude apps, API, AWS, Google Cloud, Microsoft Azure | ChatGPT Plus, Pro, Business, Enterprise; API; AWS; Azure | ChatGPT paid tiers; API |
Both flagships also offer a fast mode at a premium. Opus 5.5's fast mode costs $8 / $40, and Astra's fast mode runs up to twice as fast at twice the standard price.
The Benchmarks
A caveat first, because it matters more than usual this time. The head-to-head numbers below come from Anthropic's launch announcement, which means Anthropic chose the tests. OpenAI's launch materials emphasize different ones. Neither set is dishonest, but each shows a vendor's model at its best. Where independent testers have published results, we say so.
Head-to-Head, as Published by Anthropic
| Benchmark | Focus | Claude Opus 5.5 | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | Agentic coding in a terminal | 66.4% | 57.9% | 55.8% |
| FrontierCode v1.1 | Coding | 54.4% | 53.3% | 50.3% |
| GDPval-AA v2.1 | Work across 44 occupations (Elo) | 1846 | 1542 | 1735 |
| Humanity's Last Exam (with tools) | Expert-level questions | 67.7% | 57.2% | 65.6% |
| AutomationBench | Automation tasks | 40.0% | 41.4% | 31.4% |
| Terminal-Bench-Science 0.1 | Science tasks in a terminal | 58.7% | 64.6% | 52.6% |
Fable 5.1 is Anthropic's higher-priced model ($10 / $50), included for reference.
Taken at face value, Opus 5.5 wins four of the six shared tests, and two of those wins are large: 8.5 points on Terminal-Bench 4.0 and 10.5 points on Humanity's Last Exam. The GDPval gap of more than 300 Elo points is the most striking, because that benchmark is built to resemble real office work rather than puzzles. Astra takes AutomationBench by a hair and Terminal-Bench-Science by a clear six points.
The Fable 5.1 column tells its own story. Anthropic's new mid-priced model beats its own premium model on every test in the table.
What OpenAI Emphasizes
OpenAI's launch materials highlight gains in computer use, math, and alignment, along with security capabilities we cover below. One figure deserves attention for a reason OpenAI probably didn't intend. On ARC-AGI-3, Astra scores 99.9% using OpenAI's own test harness, but 62.7% under the standard comparison harness.
That spread is the best argument for benchmark skepticism we have seen this year. The same model lands 37 points apart depending on the scaffolding around it. When you read any benchmark, including every number in this article, ask what setup produced it.
What Independent Testers Found
- Artificial Analysis scored Opus 5.5 at max effort at 58 on its Intelligence Index (v4.3.2). Vellum reports that this put it in first place at launch, leading six of the ten component evaluations. At the default medium effort, the score drops to 51. Effort settings change results a lot, and cost rises with them.
- Sonar's code evaluation found Opus 5.5 passed 87.7% of 544 HumanEval and MBPP tasks, against 88.6% for Opus 5. That is effectively a tie, but Opus 5.5 wrote 27.5% less code to get there.
We find the Sonar result more interesting than any leaderboard position. The same pass rate with less code means less to review, less to maintain, and fewer places for bugs to hide. For teams shipping production software, that is worth more than a point or two on a leaderboard.
Independent comparisons that include both Opus 5.5 and Astra are still thin. Expect more over the coming weeks, and expect some of the vendor numbers above to shift once they arrive.
What Actually Advanced
Benchmarks are a scoreboard. The more useful question is what changed in how these models work.
Efficiency became the headline. Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5 and costs about 40% less to run on typical workloads, thanks to lower prices and fewer tokens per task. OpenAI says Astra produces shorter reasoning traces than GPT-5.6 Sol at the same effort level. After years of "bigger is better," both labs now compete on how much work you get per dollar and per second.
Long-running agentic work became practical. Anthropic's launch examples include a 680,000-line code migration finished in less than a day, and a 200,000-line codebase audited and repaired in under three hours, a job that took Opus 5 more than 20 hours. The vendor picked these anecdotes, but they describe work that was out of reach for AI models not long ago.
Both labs now restrict their own flagship models. This is the change we would underline. OpenAI rated Astra "Critical" for cybersecurity under its Preparedness Framework, the first model it has placed at that level. OpenAI defines Critical as the ability to find and exploit new vulnerabilities in hardened systems without step-by-step human guidance, and Astra discovered two zero-day vulnerabilities during pre-release testing. Its most sensitive capabilities are reserved for vetted cybersecurity defenders, and enterprise admins must switch Astra on for their workspace, because access is off by default.
Anthropic took a different route. With Opus 5.5, most cybersecurity tasks are routed to the older Claude Opus 4.8, sensitive biology requests are also redirected to older models, and unrestricted access requires enrollment in Anthropic's Cyber or Life Sciences Verification Programs. Anthropic also reports that Opus 5.5 posted the best score of any model to date on its automated behavioral audit, attempted to get around its boundaries about 85% less often than Opus 5, and is more resistant to prompt injection.
OpenAI's own safety documentation acknowledges that Astra is harder to monitor than its predecessor. Credit to OpenAI for publishing that. It is also exactly the kind of finding that makes an independent standards body feel less like a nice idea and more like a necessity. If your team does security work, the practical point is simple: neither flagship gives you its full capabilities by default.
Prices Are Falling, Just Not at the Very Top
The pricing moves this month were aggressive:
- Claude Opus 5.5: $4 / $20 per million tokens, down 20% from Opus 5's $5 / $25. Cached input reads fell 60%, from $0.50 to $0.20. Anthropic says Opus 5.5 performs at the level of Fable 5.1 on most work, and Fable 5.1 costs $10 / $50. That is roughly Fable-class output at 60% less.
- GPT-6 Sol: $2 / $10, half the price of GPT-5.6 Sol.
- GPT-6 Luna: $0.10 / $0.50, down from $0.20 / $1.20.
- GPT-6 Astra: $10 / $50, exactly the same as Fable 5.1.
The pattern is clear. The price of the single best model is holding at around $10 per million input tokens. What is collapsing is the price of last generation's frontier performance.
One caution: price per token is not cost per task. A cheaper model that reasons longer can cost more per job. OpenAI's own marketing shows how these comparisons get framed. It says Sol matches Claude Opus 5 on OSWorld 2.0 (60.5% vs. 60.3%) at about 80% lower cost per task. That comparison is against Opus 5, not Opus 5.5, which launched the same day. Test on your own workload with your own prompts before you switch anything.
For businesses, our advice is to treat your model choice as a lease, not a purchase. Build your product so the model behind it can be swapped with a configuration change, and re-run your evaluations every quarter. Teams that do this will capture every price cut. Teams that hard-code a single vendor will keep paying last year's prices.
The IPO Backdrop
All of this is happening with Wall Street watching.
Anthropic confidentially filed its S-1 registration with the SEC on June 1. Its most recent private round, a $65 billion Series H in May, valued the company at about $965 billion. It has reportedly chosen Nasdaq, and the timing has slipped. After aiming for mid-October, it is now targeting November, with late October at the earliest, according to The Wall Street Journal. Reports have floated a valuation near $2 trillion and a raise of up to $100 billion, which would make it the largest IPO on record. Anthropic is in its pre-IPO quiet period and has not confirmed a date.
OpenAI has also filed confidentially but is moving more slowly. CFO Sarah Friar told employees in August that OpenAI "will be a public company in 2027" or sooner.
We think the IPOs explain a lot about this month. Public-market investors reward efficiency and predictable margins, and both launches read as if they were written with that audience in mind: faster, cheaper per task, and heavy on enterprise use cases. The upside for everyone else is transparency. An S-1 has to disclose risks, costs, and commitments in far more detail than a launch blog post. Whatever you think of the valuations, the frontier labs are about to become much easier to scrutinize. (None of this is investment advice.)
Rivals at the Same Table: The Push for a Safety Standard
The most consequential news this month may not be a model at all.
On September 24, The Information reported that Anthropic, Google, and OpenAI are moving ahead with plans for a shared AI safety standards body. As described, it would:
- support third-party organizations that test models before release,
- define how safety and security incidents are reported,
- set voluntary safety and security commitments,
- set qualifications for independent model and lab auditors, and
- possibly run model testing of its own.
The target launch is the end of 2026 or early 2027. Its governance, membership, and relationship to the existing Frontier Model Forum, which these companies helped found in 2023, have not been settled.
The announcement capped a fast-moving summer:
- July 14: Google DeepMind CEO Demis Hassabis proposed a U.S.-led standards body modeled on FINRA, the self-regulatory organization that oversees U.S. broker-dealers.
- September 12: Anthropic CEO Dario Amodei published "We Must Pace the Frontier," arguing the industry should slow capability gains long enough for alignment work and independent verification to keep up, starting with outside evaluators embedded inside the labs. OpenAI CEO Sam Altman quickly agreed that the industry should slow down.
- September 15: OpenAI's global policy chief, Chris Lehane, confirmed on the record that OpenAI had been talking with Anthropic and Google DeepMind about safety coordination "for several weeks."
- September 23: Altman and Amodei addressed the UN Security Council. Amodei warned that "if managed poorly, I even believe AI could be a risk to humanity as a whole." He proposed narrow global agreements, such as a ban on using AI to build biological weapons, systems that let countries verify each other's commitments, and common testing standards with an incident-notification system. Speaking for the Trump administration, Michael Kratsios said the U.S. "totally reject[s] all efforts by international bodies to assert centralised control and global governance of AI."
Not everyone is applauding. Critics worry that the three largest labs could use a standards body to shut out open-source developers and smaller competitors. The FINRA comparison also cuts both ways: an industry-funded overseer has to prove it will constrain the companies that pay for it.
Our take is cautious optimism, with the emphasis on cautious. Three companies racing each other toward public listings have agreed that they should not be the only ones grading their own homework. That is real progress. A standards body is only as credible as its independence, though, and four things would convince us:
- Evaluators with real access and the authority to delay a launch.
- Published results, including the unflattering ones.
- Open-weight developers in the room as members, not as the subject of the rules.
- Incident reports the public can read.
Without those, it is a trade association with a better name.
The Verdict
Choose Claude Opus 5.5 if your work centers on software engineering, long-running agentic tasks, or knowledge work such as analysis, research synthesis, and document-heavy workflows. On the benchmarks available, it leads where most businesses spend their AI budget, and at $4 / $20 it costs 60% less per token than Astra.
Choose GPT-6 Astra if your workloads are heavy on math or science, you are building computer-use automation where it holds a narrow edge, or your organization has standardized on ChatGPT Enterprise and the switching cost isn't worth it.
Choose GPT-6 Sol or Luna if you run high volumes where cost matters more than peak capability. Sol at $2 / $10 undercuts both flagships, and Luna at $0.10 / $0.50 is priced for bulk work.
If we had to pick one model for a software team today, we would pick Opus 5.5. Based on the published numbers, it offers the most capability per dollar at the top of the market right now. We would hold that opinion loosely, though. The gaps in these tables are smaller than the headlines suggest, most of the numbers come from the vendors themselves, and the next release could reverse the order. The labs are competing hard on price and speed while starting to cooperate on safety. For those of us who build with these tools, that is the right combination.
Weighing a model switch, or planning an AI feature and want a second opinion? At Panoramic Software, we help teams choose, integrate, and evaluate AI models for real products. Let's talk about your project.
Sources
- Anthropic: Introducing Claude Opus 5.5
- OpenAI: GPT-6 Astra: A new generation of intelligence and Safety overview: GPT-6 Astra
- VentureBeat: Anthropic releases Claude Opus 5.5
- MarkTechPost: OpenAI releases GPT-6 Sol and Luna
- AlphaCorp AI: GPT-6 Astra launch: benchmarks and pricing
- InfoQ: GPT-6 Astra is the first model OpenAI classifies as Critical for cybersecurity
- WinBuzzer: Claude Opus 5.5, GPT-6 Sol and Luna launch at lower prices
- Artificial Analysis: Claude Opus 5.5 release intelligence
- Vellum: Claude Opus 5.5 benchmarks explained
- Sonar: Claude Opus 5.5: an evaluation
- CNBC: Anthropic confidentially files IPO prospectus and OpenAI "will be a public company in 2027" or sooner
- Forbes: Anthropic IPO slips to November
- PYMNTS: OpenAI, Google and Anthropic join forces to set AI safety standards
- Dario Amodei: We Must Pace the Frontier
- Al Jazeera: AI leaders tell UN the industry needs global regulation
- CNN: Altman and Amodei urge UN Security Council to adopt international AI standards
