Anthropic released Claude Opus 5 on Friday, positioning it not as its smartest model but as its most economically compelling one — a deliberate pivot in an industry where enterprise AI budgets are now treated as serious operational costs rather than R&D experiments.

The model is priced at $5 per million input tokens and $25 per million output tokens, unchanged from its predecessor Opus 4.8. It immediately becomes the default on Claude Max (Anthropic's premium consumer tier) and the strongest model available on Claude Pro.

The Capability vs. Cost Argument

Anthropics's positioning is explicit: Opus 5 delivers nearly all the intelligence of the top-tier Claude Fable 5 at roughly half the cost. Rather than chasing benchmark supremacy, the company is arguing that most enterprise AI work sits in a "middle band" of difficulty — complex enough to need near-frontier intelligence, but frequent enough that per-token economics dominate the decision.

"Opus 5 as your daily driver, the model you hand complex work to and review when it's done. Fable 5 for your most ambitious work, the days-long autonomous projects nothing could take on before. Sonnet 5 for work you run at scale, where speed and cost per call decide what ships."

This stratification is Anthropic's clearest articulation yet of a tiered model lineup where each product occupies a distinct economic niche.

Benchmark Numbers — and Honest Caveats

The headline figures are striking:

  • Frontier-Bench v0.1 (agentic terminal coding): Opus 5 scores 43.3%, up from Opus 4.8's 18.7% — and above Fable 5's 33.7%, at lower cost
  • ARC-AGI 3 (novel problem-solving): Anthropic claims three times the score of the next best model
  • OSWorld 2.0 (computer-use): surpasses Fable 5's best result at just over one-third the cost

Notably, Anthropic doesn't bury the gaps. The company acknowledges Opus 5 trails a competing model (referred to as Mythos 5) on cybersecurity and biology tasks, and an OpenAI-family model still leads on at least one agentic coding benchmark.

The more revealing caveat: Opus 5 wins on bounded evaluations with defined outcomes. Fable 5 retains an edge on long-horizon tasks — jobs requiring coherence across many steps over hours or days. That bounded-vs.-autonomous axis may become the defining dimension of model differentiation through 2026, as standard benchmarks saturate.

Token Efficiency Is the Real Battleground

Enterprise customers offered unusually specific validation of the efficiency claims:

  • Harvey (legal AI): Opus 5 matched Opus 4.8's maximum-reasoning mode while generating 26% fewer tokens on average
  • Fundamental Research Lab: On financial-modeling tasks, nine percentage points higher accuracy using roughly one-third fewer turns and 60% less time
  • Zapier: Opus 5 topped its AutomationBench leaderboard, completing a full churn-prevention workflow end-to-end — "Previous models didn't pass; Opus 5 hit 100%"
  • Cognition (maker of the Devin coding agent): "Approaches Fable-level performance at half the cost," with particular strength in debugging and root-cause analysis

Anthropics also ships the model with an adjustable "effort" setting, letting customers explicitly trade intelligence for speed and token savings depending on the task.

This matters commercially. According to a February 2026 analysis by Contrary Research, Claude held roughly 40% of the enterprise LLM market by usage as of late 2025. Claude Code alone had reached approximately $1 billion in annualized revenue. For a company whose customers pay by the token, efficiency isn't a feature — it's the core value proposition.

Self-Verification as the Missing Link for Deployment

Beyond scores, Anthropic is emphasizing a behavioral property: Opus 5 verifies its own work and iterates until it succeeds.

The company shared testing vignettes that illustrate this well:

  • Asked to reconstruct a machine part as a 3D CAD model from a drawing it couldn't view, Opus 5 wrote its own computer vision pipeline to extract geometry from raw pixels — and kept iterating while no competing model solved the task in five attempts
  • Given a real bug in a popular open-source package manager, the model found the root cause and fixed an edge case the community's own patch had missed
  • A trading firm engineer used Opus 5 to build a market data feed in a single session; finding no live feed to validate against, the model built its own test harness

One Stripe staff engineer described giving the model a "chief-of-staff role over my dev environments" for a weekend: "It built its own monitor, drove each box, and pulled me in only for the judgment calls."

This self-verification loop closes the gap between demo and deployable system. Most hidden cost in enterprise AI today is human review — engineers auditing model output. A model that reliably checks its own work compresses that overhead, which is why customers keep citing fewer turns and less time rather than higher raw scores. For startup founders and enterprise buyers alike, that's the metric that actually moves a business case.

What This Means for Competitors

Google, OpenAI, and others are navigating the same tradeoff space — pushing capability at the frontier while trying to compress costs at the tier below it. Anthropic's move with Opus 5 sharpens the competitive argument that the mid-tier model slot, not the frontier flagship, is where enterprise revenue is actually won or lost.