Claude Opus 4.8 Fast
About this model
Claude Opus 4.8 Fast is the latency-optimized configuration of Claude Opus 4.8, Anthropic's most capable generally available model, released on May 28, 2026. Rather than a separate model, fast mode serves the same Opus 4.8 weights at higher throughput: setting speed to "fast" yields up to 2.5× more output tokens per second at premium pricing, offered initially as a research preview on the Claude API. It retains the full 1M-token context window that runs by default on the Claude API, Amazon Bedrock, and Vertex AI, plus up to 128k max output tokens and adaptive thinking.
Compared with the prior generation, Claude Opus 4.7 Fast, the underlying 4.8 model targets behavioral gains that Anthropic attributes to long-horizon agentic coding, including better long-context handling, fewer compactions, and improved compaction recovery. Anthropic also reports better tool triggering — the model is less likely to skip a required tool call, an issue some users flagged on Opus 4.7 — and improved honesty, with reduced tendency to overclaim progress.
A notable practical change: Anthropic states that fast mode for Opus 4.8 is three times cheaper than fast mode was on previous Opus models, narrowing the cost gap with the standard tier.
This is the newest entry in the Opus fast lineage. It sits alongside the broader Claude family, including Claude Fable 5, which Anthropic describes as its most capable model in Claude Code for tasks larger than a single sitting.
This About section is AI-generated from public sources (Claude Opus 4.8), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 3d ago