Claude Opus 4.8 · API · High effort (default)
The expensive architect — pays for hard repos, long agents, and fewer retries.
Flagship coding ceiling when constraint-heavy, multi-hour, or autonomous work justifies $5/$25.
Overall mean of attribute public scores after peer min-max (40–95 band within scored launch peer set).
- Endpoint
- claude-opus-4-8
- Product
- Anthropic API
- Price
- $5 / $25 per MTok
- Context
- 1000k tokens
Flagship Opus 4.8 API, high effort, 1M context, $5/$25 — expensive coding ceiling.
Six stats · peer public
Peer public (main number / bar) ranks this Build inside the current scored peer set (≈40–95). Raw is the absolute evidence-fusion aggregate before min-max. A low peer score can still be a strong absolute model — see methodology.
When to choose
- Hard repository migration and multi-file refactors
- Autonomous coding agents with named harnesses
- High-stakes reasoning and verification loops
- Everyday chat and low-stakes drafts (use Sonnet/Flash)
- Strict budget / high-volume token farms
- Latency-first interactive sprinters
The job is hard enough that a wrong answer costs more than Opus tokens.
Task proficiencies
Signature strengths and honest flaws
Sustains multi-hour, multi-file autonomous coding runs without losing the thread — holds architecture, earlier decisions, and constraints across long agent loops better than lighter Builds.
- Named-harness coding agents on hard repository migrations
- Multi-hour refactors and long-horizon execution
Tends to self-check and catch its own errors in high-stakes reasoning before committing, reducing wrong answers on jobs where a mistake costs more than the extra tokens.
- High-stakes reasoning and verification loops
- Constraint-heavy work where a wrong edit is expensive
The most expensive Build in the scored set ($5/$25 per MTok). Token spend climbs fast on high-volume, high-frequency, or long autonomous runs.
- Autonomous agent loops without output budgets
- High-volume or repeated full-file rewrites
Workaround · Reserve Opus for genuinely hard sessions; route bulk, simple, or draft work to Sonnet or a cheaper frontier rival. Set explicit token/time budgets on agents.
Over-deliberates low-stakes chat and quick drafts, adding latency and cost where a sprinter Build would finish the job just as well.
- Everyday chat and low-stakes drafts
- Latency-first interactive tasks
Workaround · Use lower-effort or lighter Builds (Sonnet, Flash) for everyday interactive work; escalate to Opus only when the job is hard enough to justify it.
Traceability
Every factual claim cites an evidence row. Tap a chip to see the benchmark, exact config, source, and caveats behind the number.
Engine pass meth_v0_1_1 (peer min-max vs scored launch peer set). Codex overall 66.9 (B). Attribute public scores map raw aggregates into a 40–95 band within the peer set (methodology min-max). INT uses hierarchical fusion (GPQA primary / HLE secondary / capped bleed).