Claude Opus 4.6
Anthropic's December 2025 Opus refresh, with the headline addition of a 1-million-token context variant. The agentic headline: sustained autonomous coding sessions extending beyond 12 hours uninterrupted (vs Opus 4's ~7 hour benchmark from May 2025). The model behind most "AI did multi-day project autonomously" demos in late 2025.
How are Intelligence, Speed & Cost bucketed?
- Top 1%≤ 1%
- Top 5%≤ 5%
- Top 10%≤ 10%
- Good≤ 25%
- Medium≤ 50%
- Below avg> 50%
- Top 1%≥ 345 tok/s
- Top 5%≥ 237 tok/s
- Top 10%≥ 196 tok/s
- Good≥ 146 tok/s
- Medium≥ 90 tok/s
- Slow< 90 tok/s
- Freeopen weights · self-host
- Low< $1 / M out
- Moderate$1–5 / M out
- High≥ $5 / M out
Why it matters
Opus 4.6 is the model that this entire codebase is likely being authored by — its capability profile is what made "build a content site with 50+ MDX nodes plus dual orientation graph + light theme + Civ-style layout in one session" a tractable conversation rather than a multi-week project.
Core Capabilities
Context Window
Availability
Pricing Model
Capability / Performance
Where this model sits relative to the middle 60% of models in the tree. All scores are 0–10 (higher is better).
What it feels like
- First Opus with 1M-token context window — the long-context gap to Gemini 2.5 Pro finally closes
- Tops Terminal-Bench 2.0 (65.4%), OSWorld (72.7%), BrowseComp (84.0%), Finance Agent (60.7%)
- 53.1% on Humanity's Last Exam with tools — beats every other frontier model at release
- ARC-AGI-2 score of 68.8% — biggest single-jump on novel-problem-solving in any model so far
- Adaptive thinking + conversation compaction make multi-hour sessions much more stable
- SWE-bench Verified 80.8% (vs 80.9% on Opus 4.5) — coding plateaued; the gains are in agents and reasoning
Best use cases
- Long-horizon agent loops with full 1M-token codebases or document corpora
- Web-research and browse-style agents (BrowseComp leader)
- Financial analysis, forecasting, and structured-data agent flows
- Office-style multi-step tasks (GDPVal-AA leader at 1606 Elo)
Tools to try
Not ideal for
- Latency-sensitive interactive chat — adaptive thinking adds time
- Pure SWE-bench coding — Opus 4.5 at the same score is cheaper
- Workloads where 200K context is plenty — paying for 1M isn't justified
Model Evolution
claude-opus is Anthropic's language model family.