Anthropic released a new AI model today called Claude Opus 5, and it matters for one plain reason: it delivers close to top-tier intelligence at a much lower price than the models that used to be required for serious automation work.
If you run a business and you've been told that reliable AI automation needs "the expensive model," this launch is worth five minutes of your attention.
What actually happened
Anthropic — the company behind the Claude family of AI models — announced that Claude Opus 5 is available today across its platforms, including the Claude API that automation systems are built on. In Anthropic's own words, it is a "thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price."
Fable 5 is Anthropic's flagship, top-of-the-line model. Opus 5 is the tier below it. The news is how small the gap between those tiers has become, and how large the price difference remains.
The pricing, from Anthropic's announcement: $5 per million input tokens and $25 per million output tokens — unchanged from the previous Opus 4.8. (A "token" is a small chunk of text; models are billed by how much text they read and write.) There's also a new fast mode that runs at around 2.5x the default speed for twice the base price, for work where response time matters more than cost.
The plain-language version of the benchmarks
Anthropic published evaluation results alongside the launch. Three stand out for anyone whose interest in AI is "will it do the work correctly":
Automation tasks. On Zapier's AutomationBench — a test of whether a model can correctly carry out multi-step business automations — Anthropic reports Opus 5 achieved roughly 1.5x the pass rate of the next-best model for the same cost per task, and per Zapier, it hit 100% on an account-health workbook task that previous models didn't pass. Automation is exactly the category where errors are expensive, because an agent that's wrong 10% of the time creates cleanup work instead of saving it.
Computer use. On OSWorld 2.0, which measures a model's ability to operate software the way a person does, Anthropic reports Opus 5 surpassed its own flagship Fable 5 at just over a third of the cost.
Coding. Anthropic reports state-of-the-art results on Frontier-Bench, more than doubling Opus 4.8's performance at a lower cost per task, and — at max effort — coming within 0.5% of Fable 5 on CursorBench 3.2 at half the cost per task.
Zapier CEO Wade Foster put it this way in the announcement: "Claude Opus 5 topped Zapier's AutomationBench leaderboard without spending more tokens than prior Claude models."
One honest caveat from Anthropic's own materials: Opus 5 still trails a competing model (Mythos 5) on some cybersecurity and biology research tasks. No single model wins everything.
What this means if you're not a tech company
Two concrete examples of where the cost-per-quality shift lands.
Property management. The AI workloads in property management are high-volume and repetitive: triaging maintenance requests, answering tenant calls after hours, chasing rent-related paperwork, summarizing inspection reports. High-volume is exactly where per-token price dominates the economics. When near-flagship reasoning becomes available at the mid-tier price, workloads that were previously "too expensive to run on the good model" — like reading every incoming maintenance request and routing it with correct urgency — become viable to run on a model that rarely makes routing mistakes.
Transportation and logistics. Dispatch and freight operations live on messy, unstructured inputs: rate confirmations, bills of lading, driver messages, exception emails. This is reasoning-heavy work, and until now the trade-off was real — cheaper models misread edge cases, flagship models cost too much to run on every document. A model that Anthropic reports doubles its predecessor's performance on hard reasoning benchmarks, at the same per-token price as that predecessor, shifts the answer on "can we afford to run this on everything."
HVAC and home services. The after-hours phone problem is the classic one: a compressor fails at 9pm, the caller gets voicemail, and the job goes to whoever answers. AI phone agents already handle this — but the difference between a model that merely takes a message and one that reasons well is the difference between "someone will call you back" and an agent that correctly distinguishes a no-heat emergency from a routine maintenance booking, quotes the right service window, and escalates the genuine emergencies to the on-call tech. Reasoning quality is what makes that triage trustworthy, and reasoning quality at mid-tier prices is exactly what today's release moved.
Questions worth asking your automation provider this week
You don't need to evaluate models yourself — that's your provider's job. But three questions will tell you quickly whether they're doing it. Which model runs my workflows today, and when was that choice last revisited? What would change — in accuracy or in cost — if it moved to a newer model at the same price tier? And is there a workload we previously ruled out on cost that's worth re-scoping now? A provider who can answer those crisply is paying attention. One who can't is running your business on last year's assumptions.
The takeaway
The pattern across the last two years of AI releases is consistent: capability moves down-market faster than most businesses re-evaluate. Systems that were quoted or scoped six months ago were priced against a different capability-per-dollar curve than the one that exists as of this morning.
You don't need to become an AI expert to act on that. The useful move is simpler: if you have automation running — or a proposal sitting in your inbox — it's a reasonable week to ask what model it runs on, and whether today's release changes what's possible at your budget.




