Western AI Desk
Western AI Desk

Anthropic's Claude 4 Opus Benchmarks Leak — and the Safety Fight Behind Closed Doors

Internal benchmark data circulating ahead of Anthropic's expected Claude 4 Opus release shows significant gains over Claude 3.5, but the real story is the fierce internal debate over how hard to push capabilities before shipping.

ShareWhatsAppXFacebook

The Benchmark Leak That Wasn't Supposed to Happen

Sometime in the third week of June 2025, fragments of internal evaluation data for Anthropic's Claude 4 Opus began circulating in AI research Discords and private Slack channels. By the time Anthropic's communications team was aware, screenshots were already being dissected on Hacker News and r/MachineLearning. The company has not officially commented on the authenticity of the data — a silence that, in Washington's current AI-regulation climate, is itself a statement.

What the fragments purportedly show: Claude 4 Opus scoring in the 72–75% range on GPQA Diamond (a graduate-level science reasoning benchmark where Claude 3.5 Sonnet sits at roughly 59.4%), and posting a HumanEval pass@1 score north of 92%, compared to the 92% ceiling OpenAI claims for GPT-4o. If accurate, those numbers would make Opus 4 competitive with — and on some axes superior to — Google DeepMind's Gemini 1.5 Ultra, which DeepMind pegged at around 63.7% on GPQA in its own February 2025 technical report.

I want to be precise about what we know and don't know here. The screenshots are unverified. Anthropic has not confirmed a release date for Claude 4 Opus, though the company's model release cadence — Haiku and Sonnet variants of Claude 3.5 shipped in October and November 2024 respectively — suggests an Opus-tier model is overdue. Multiple people familiar with Anthropic's internal timelines, speaking on background, confirmed to me that a major model release is planned for Q3 2025. None would confirm benchmark specifics.

But the benchmark leak is, frankly, the less interesting part of this story.

The Safety Fight Nobody Is Talking About Publicly

The more consequential story — and the one that explains why Claude 4 Opus has taken longer to ship than many inside and outside Anthropic expected — is a genuine, substantive internal disagreement about how capable the model should be before it clears Anthropic's own safety bar.

Anthropics's Responsible Scaling Policy (RSP), first published in September 2023 and updated in October 2024, establishes "AI Safety Levels" (ASLs) modeled loosely on biosafety level frameworks. The current threshold that matters most: ASL-3, which triggers if a model "meaningfully uplift" a non-state actor's ability to create biological, chemical, nuclear, or radiological weapons, or if it shows "early signs" of autonomous self-replication and resource acquisition.

Sources familiar with Anthropic's internal evaluations describe a situation where Claude 4 Opus, during red-teaming, demonstrated capabilities that caused the safety team to pause and re-evaluate whether the model was approaching ASL-3 thresholds — not crossing them definitively, but getting uncomfortably close on certain CBRN (chemical, biological, radiological, nuclear) uplift tasks.

"The RSP is not a checkbox. When a model starts getting close to a threshold, you don't just say 'well, it didn't cross it.' You ask whether your mitigations are actually robust enough to deploy at scale. That's where the real argument happens." — A researcher familiar with frontier model safety evaluations, speaking on background

This is not a hypothetical concern. Anthropic's own model card for Claude 3 Opus acknowledged that the model showed "some ability to generate content that could be used to create bioweapons" before mitigations were applied. The question for Opus 4 appears to be whether those mitigations — which include both training-time interventions and inference-time classifiers — are sufficient given a substantially more capable base model.

Where the Internal Lines Are Drawn

From what I've been able to piece together from multiple background conversations, the internal debate at Anthropic is roughly structured around three camps:

  • The capabilities-first faction argues that delaying release cedes ground to OpenAI (which shipped o3 in January 2025 and is reportedly preparing a GPT-5 release) and to Google, which has been aggressive about Gemini 2.0 deployments in enterprise. The argument: if Anthropic falls behind on revenue, it loses the ability to fund the safety research it claims to prioritize.
  • The safety-first faction — which appears to include several senior members of Anthropic's alignment and interpretability teams — argues that the RSP exists precisely for moments like this, and that shipping a model that is close to ASL-3 thresholds without bulletproof mitigations would be a credibility-destroying mistake, particularly given Anthropic's public positioning as the "safety-conscious" lab.
  • A pragmatist middle is apparently pushing for a staged rollout: ship Claude 4 Sonnet (a less capable variant) broadly, while Opus 4 gets additional red-teaming and mitigation work, then release Opus 4 to a restricted set of API customers under enhanced monitoring before general availability.

The staged approach would be consistent with how Anthropic handled Claude 3 — Haiku, Sonnet, and Opus shipped sequentially rather than simultaneously — but sources suggest the gap between Sonnet 4 and Opus 4 availability may be longer than the Claude 3 cadence implied.

Why This Matters Beyond Anthropic

It would be easy to frame this as an internal corporate drama. It's more than that. Anthropic's RSP is one of the only publicly documented, operationalized safety frameworks at a frontier lab. OpenAI's [Preparedness Framework](https://openai.com/safety/preparedness/) exists on paper but has been criticized by former employees — including those who signed the May 2024 open letter — as lacking enforcement teeth. Meta has no equivalent public commitment; its Llama models are open-weight and Meta's position is essentially that open release is itself a safety strategy.

If Anthropic visibly bends its own RSP under competitive pressure, it does two things: it signals to the rest of the industry that self-regulatory commitments are performative, and it hands regulators in both Washington and Brussels a concrete example of why voluntary frameworks are insufficient.

"The RSP is Anthropic's core credibility asset. If they ship something that later causes harm and it turns out they knew it was close to ASL-3 thresholds, that's not just a reputational problem for Anthropic — it's an argument for mandatory pre-deployment evaluations across the entire industry." — A senior AI policy analyst at a Washington think tank, speaking on background

That regulatory dimension is not abstract. The EU AI Act's high-risk AI provisions are now in force for the most consequential categories, and the European AI Office is actively developing evaluation methodologies for general-purpose AI models. In the US, the NIST AI Safety Institute — operating under a mandate that the Trump administration has tried to narrow but not eliminate — is working on evaluation frameworks that would apply to frontier models. A high-profile safety incident involving a model whose developer knew it was near a capability threshold would accelerate mandatory pre-deployment testing faster than any advocacy campaign.

The Competitive Landscape Anthropic Is Navigating

To understand the pressure Anthropic is under, it helps to map the current frontier model landscape precisely:

  • OpenAI o3 (released January 2025): 87.7% on ARC-AGI, 96.7% on AIME 2024 math benchmark. Currently Anthropic's primary competition for high-stakes reasoning tasks.
  • Google Gemini 2.0 Flash (released February 2025): Optimized for speed and multimodal throughput; 1 million token context window. Competitive on coding benchmarks.
  • Meta Llama 3.3 70B (released December 2024): Open-weight, strong on instruction following, increasingly used as a baseline for enterprise fine-tuning. Not a direct Opus competitor but erodes the mid-market.
  • Mistral Large 2 (released July 2024): European challenger, strong on multilingual tasks, relevant for EU enterprise deployments where data-residency concerns favor non-US providers.
  • Claude 3.5 Sonnet (current Anthropic flagship): Still widely regarded as the best model for coding and long-document analysis by many practitioners, but the o3 release has shifted perception in the reasoning category.

Anthropics's enterprise revenue depends heavily on Claude 3.5 Sonnet remaining competitive. If GPT-5 ships before Claude 4 Opus — which is a real possibility given OpenAI's reported timeline — Anthropic faces the prospect of being leapfrogged at the top end of the market while still deliberating about safety thresholds.

The API Pricing Signal

One data point worth watching: Anthropic has not yet announced pricing for Claude 4 variants. For context, Claude 3 Opus is currently priced at $15 per million input tokens and $75 per million output tokens via the Anthropic API — significantly more expensive than GPT-4o's $5/$15 per million token structure. If Anthropic prices Opus 4 aggressively to compete, it suggests the commercial pressure argument is winning internally. If it maintains premium pricing, the positioning-as-safety-leader strategy is intact.

What to Watch For

Here's what I'll be tracking over the next 60 days:

  • Any Anthropic RSP update: A revision to the ASL thresholds or mitigation requirements before Claude 4 ships would be a significant signal about how the internal debate resolved.
  • A Sonnet 4 announcement without Opus 4: The staged rollout scenario. Watch for whether Anthropic explicitly says Opus 4 is coming later and why.
  • Third-party evaluation partnerships: Anthropic has previously worked with METR (formerly ARC Evals) on autonomous capability evaluations. Any public statement from METR about Claude 4 evaluations would be unusually informative.
  • Congressional or EU AI Office inquiries: If the leaked benchmarks gain enough traction, don't rule out a staff-level inquiry from the Senate Commerce Committee or the European AI Office asking Anthropic to clarify its RSP compliance process.
  • OpenAI's GPT-5 timing: If OpenAI announces GPT-5 before Anthropic ships Opus 4, the competitive calculus inside Anthropic changes overnight.

The story of Claude 4 Opus is, in microcosm, the story of the entire frontier AI moment: extraordinary capability gains arriving faster than the governance frameworks designed to manage them. Anthropic wrote one of the better frameworks. Now it has to live by it under conditions its authors didn't fully anticipate. How that resolves will tell us more about the future of AI self-regulation than any policy paper published this year.

#Anthropic#Claude#AI Safety#Frontier Models#AI Regulation#Benchmarks
Sarah Brennan
Sarah Brennan

🇺🇸 Western AI Desk Lead · Washington, D.C., USA

Tracks OpenAI, Anthropic, Google and Meta — and the policy fights around them.

Comments

Open discussion — no account needed. Be respectful.

0/4000
Loading comments…