Updated October 7, 2026 with Claude Haiku 5.5
Token Estimator – paste text to estimate tokens and cost
Model Comparison – see costs across all models
| Model | Per Request | 10K Requests |
|---|
Calculating Claude API Costs
Anthropic charges per million tokens, with separate rates for input and output. Total Claude API costs depend on which model you select, how much of your context repeats across requests, and whether your workload can run asynchronously. This page covers current pricing as of October 7, 2026, including Claude Haiku 5.5, Claude Sonnet 5.5 and Claude Opus 5.5, the features that reduce spend, and the considerations that matter when choosing a model.
Paying for API Access
The Claude API is pay as you go and billed on usage. There is no subscription required to call it, and no monthly minimum. An API key itself costs nothing to create, and holding one costs nothing; you are billed only for the tokens your requests actually consume. New organizations receive a small amount of free credit to test with, after which requests need a funded balance.
API billing is separate from a Claude subscription, and API spend never counts against a plan’s usage limits. Since October 7, 2026, Max and Team plans include a monthly API credit: $100 on Max 5x, $200 on Max 20x, and on Team $20 per Standard seat and $100 per Premium seat, pooled up to $500. Free, Pro and Enterprise plans are not eligible. Anthropic is rolling the credit out over a few days, and a plan has to be active for seven days before it can be claimed.
The credit goes to one Claude Console organization and covers the Claude API, including batch, the Console Playground, Managed Agents and the Agent SDK. It does not cover interactive Claude Code or Claude on Amazon Bedrock, Google Cloud or Microsoft Foundry. Unused credit expires at the end of each billing cycle. When it runs out, requests stop unless the organization has bought credits or turned on auto-reload, so the credit on its own cannot run up a bill. For a Max or Team subscriber, subtract the monthly credit from the calculator’s monthly total to see what the organization would pay on top of it.
Current Claude Models
Anthropic released Claude Haiku 5.5 on October 7, 2026. It is the only current Claude model priced by prompt length: $0.10 input and $0.50 output per million tokens for prompts up to 100,000 tokens, and five times that, $0.50 / $2.50, for longer prompts. Even above 100,000 tokens it costs half of Haiku 4.5 per token. The same day Anthropic halved Claude Sonnet 5.5’s cache reads to $0.10 per million tokens; its $2 / $10 base rates are unchanged. Claude Opus 5.5 at $4 / $20 is still the model Anthropic suggests starting with for most workloads, and Fable 5.1 at $10 / $50 is for demanding reasoning and long-horizon agentic work, or for when Opus 5.5 at higher effort still falls short. Claude Mythos 5.1 is the same model as Fable 5.1 with more permissive safeguards, available only through Anthropic’s trusted access programs. Claude Sonnet 4.5 was deprecated on September 30 and retires on November 30, 2026.
Released October 7, 2026 for high-volume, latency-sensitive work: classification, extraction, routing, summaries, and subagent tasks under a Sonnet or Opus orchestrator. Use the API identifier claude-haiku-5-5. On prompts up to 100,000 tokens, cache writes are $0.125 (5-minute) and $0.20 (1-hour), cache reads $0.01 and batch $0.05 / $0.25. Over 100,000 tokens every line is five times higher: $0.50 input, $2.50 output, $0.625 and $1 cache writes, $0.05 cache reads and batch $0.25 / $1.25. Anthropic has not said whether cached tokens count toward the 100,000, so budget as if they do. The minimum cacheable prompt is 512 tokens, there is no fast mode, and default effort is medium. Available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS.
Released September 28, 2026 and described by Anthropic as its best combination of speed and intelligence, strongest at well-scoped everyday tasks, bug fixes, and documents, slides and spreadsheets. Use the API identifier claude-sonnet-5-5. Cache writes are $2.50 (5-minute) and $4 (1-hour) and batch is $1 / $5, the same as Sonnet 5. Since October 7, 2026, cache reads cost $0.10 per million tokens, half of Sonnet 5’s $0.20; Anthropic estimates this makes Sonnet 5.5 about 20% cheaper on most agentic work, where cache reads are a large share of the tokens. The minimum cacheable prompt drops to 512 tokens from Sonnet 5’s 1,024. There is no fast mode. Default effort is high on the Claude API and medium in Claude Code and the Claude apps. Available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS.
Released September 22, 2026, and the model Anthropic suggests starting with for most workloads. Use the API identifier claude-opus-5-5. Every rate is below Opus 5: cache writes are $5 (5-minute) and $8 (1-hour), cache reads $0.20, batch $2 / $10 and fast mode $8 / $40. The minimum cacheable prompt stays at 512 tokens. Thinking cannot be switched off and the default effort is medium rather than high, so effort is the control for how many tokens a request spends. It is the default model in Claude Code on Pro, Max, Team and Enterprise plans. Available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS.
Anthropic’s model for demanding reasoning and long-horizon agentic work, released September 1, 2026. Use the API identifier claude-fable-5-1. Base pricing matches Fable 5, but cache reads cost $0.25 per million tokens instead of $1.00, so any workload that reuses a long prefix is cheaper on 5.1 at the same headline rate. Cache writes are unchanged at $12.50 (5-minute) and $20 (1-hour), and the minimum cacheable prompt is 512 tokens. Default effort is high; low and medium effort match or beat Fable 5 at a lower cost per task. Available on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry.
Identical to Fable 5.1 in weights and price, including the $0.25 cache read, with more permissive safeguards for cybersecurity and life-sciences work. Access is through the Cyber Verification Program and the Life Sciences Verification Program, currently limited to a set of US organizations. Claude Mythos 5, the counterpart to Fable 5, remains on the same limited-access terms at Fable 5 rates.
The previous Opus generation, priced identically to Opus 4.8 and above Opus 5.5 on every line. Use the API identifier claude-opus-5. It still accepts two request shapes that Opus 5.5 rejects: forced tool use and, at effort high or below, thinking disabled. Integrations that rely on either need changing before they move to Opus 5.5. Its minimum cacheable prompt is 512 tokens, and the effort setting from low to max controls how many tokens it spends.
The previous Fable generation, superseded by Fable 5.1 on September 1, 2026 at the same $10 / $50 base rate. Use the API identifier claude-fable-5. The cost difference between the two is entirely in cache reads: $1.00 per million tokens here against $0.25 on Fable 5.1. Fable 5 ships with safety classifiers that can decline certain requests; whether a refusal that arrives before any output is billed depends on its category, and the API can automatically retry on another Claude model via the fallbacks parameter.
An earlier Opus generation at the same $5 / $25 as Opus 5. Use the API identifier claude-opus-4-8. Fast mode is available for Opus 4.8 at 2x standard token pricing when speed is worth the premium.
The previous Sonnet generation, released June 30, 2026, at the same $2 / $10 base rates as Sonnet 5.5, but its cache reads stay at $0.20 per million tokens against $0.10 on Sonnet 5.5. Use the API identifier claude-sonnet-5. It still accepts request shapes that Sonnet 5.5 rejects: thinking disabled, forced tool use, and on the Claude API and Google Cloud the earlier computer use tool, computer_20251124. Its minimum cacheable prompt is 1,024 tokens. Sonnet 5 uses the same tokenizer as Sonnet 5.5, so the same text produces the same token count on both.
Superseded by Sonnet 5 and Sonnet 5.5, which both cost less at $2 / $10. Sonnet 4.6 stays available at $3 / $15 for workloads pinned to its behavior or to the older tokenizer.
The previous Haiku generation, still active, with retirement not sooner than October 15, 2026. Use the API identifier claude-haiku-4-5. Haiku 5.5 costs less on every line at any prompt length, so keep Haiku 4.5 only where a workload depends on something Haiku 5.5 rejects: assistant prefill, non-default temperature, top_p or top_k, manual thinking budgets, or Priority Tier. Its minimum cacheable prompt is 4,096 tokens, and it does not support US-only data residency.
Full Claude API Pricing Table
Standard prices per million tokens, ordered by model family: Fable, Opus, Sonnet, then Haiku. Prompt caching and batch discounts apply at standard rates unless a feature says otherwise. Retired models are included only for historical comparison. Claude Haiku 5.5 is the one current model priced by prompt length, so its row gives both prices.
| Model | Input | Output | Status |
|---|---|---|---|
| Fable family – frontier capability | |||
| Claude Fable 5.1 | $10.00 | $50.00 | Active; 1M context; cache reads $0.25 |
| Claude Mythos 5.1 | $10.00 | $50.00 | Limited availability; same model as Fable 5.1; cache reads $0.25 |
| Claude Fable 5 | $10.00 | $50.00 | Active; previous Fable generation; cache reads $1.00 |
| Claude Mythos 5 | $10.00 | $50.00 | Limited availability via Project Glasswing; 1M context |
| Opus family – flagship workhorse | |||
| Claude Opus 5.5 | $4.00 | $20.00 | Active; newest Opus; 1M context; cache reads $0.20; fast mode 2x |
| Claude Opus 5 | $5.00 | $25.00 | Active; previous Opus; 1M context; fast mode 2x |
| Claude Opus 4.8 | $5.00 | $25.00 | Active; 1M context; fast mode 2x |
| Claude Opus 4.7 | $5.00 | $25.00 | Active; 1M context; fast mode withdrawn 24 Jul 2026 |
| Claude Opus 4.6 | $5.00 | $25.00 | Active; 1M context; fast mode discontinued 29 Jun 2026 |
| Claude Opus 4.5 | $5.00 | $25.00 | Active; 200K context |
| Claude Opus 4.1 | $15.00 | $75.00 | Retired Aug 5, 2026; still on Bedrock and Google Cloud |
| Claude Opus 4 | $15.00 | $75.00 | Retired June 15, 2026; still on Google Cloud |
| Sonnet family – balanced | |||
| Claude Sonnet 5.5 | $2.00 | $10.00 | Active; newest Sonnet; 1M context; cache reads $0.10 since Oct 7, 2026 |
| Claude Sonnet 5 | $2.00 | $10.00 | Active; previous Sonnet; 1M context; cache reads $0.20 |
| Claude Sonnet 4.6 | $3.00 | $15.00 | Active; 1M context |
| Claude Sonnet 4.5 | $3.00 | $15.00 | Deprecated Sep 30, 2026; retires Nov 30, 2026 |
| Claude Sonnet 4 | $3.00 | $15.00 | Retired June 15, 2026; still on Bedrock and Google Cloud |
| Claude 3.7 Sonnet | $3.00 | $15.00 | Retired on Claude API; historical only |
| Haiku family – fastest and lowest cost | |||
| Claude Haiku 5.5 | $0.10, or $0.50 over 100K | $0.50, or $2.50 over 100K | Active; newest Haiku; 1M context; cache reads $0.01, or $0.05 over 100K |
| Claude Haiku 4.5 | $1.00 | $5.00 | Active; previous Haiku; 200K context |
| Claude Haiku 3.5 | $0.80 | $4.00 | Retired on Claude API; historical only |
| Claude Haiku 3 | $0.25 | $1.25 | Retired on Claude API; historical only |
Claude Mythos 5 shares Fable 5’s capabilities and pricing but is offered only in limited release to approved customers through Anthropic’s Project Glasswing, so Fable 5 is the generally available model in that class. Both carry 30-day data retention and are not available under zero data retention agreements, which matters for compliance-sensitive workloads. Opus 4.1 and Opus 4 cost $15 / $75, nearly four times the Opus 5.5 rate, and both are retired on the Claude API; they remain only on Amazon Bedrock or Google Cloud. Moving that traffic to Opus 5.5 cuts the token bill by about three quarters. Haiku 3, the $0.25 / $1.25 model, was retired on April 20, 2026. Haiku 5.5 is now the cheapest Claude model on the Claude API, and on prompts up to 100,000 tokens it costs less than Haiku 3 did. Sonnet 4.5 retires on November 30, 2026; Anthropic names Sonnet 5.5 as its replacement, which costs a third less on input and output.
Opus Fast Mode Costs
Fast mode runs Opus up to 2.5x faster for 2x standard token pricing: $8 per million input tokens and $40 per million output tokens on Opus 5.5, or $10 and $50 on Opus 5 and Opus 4.8. It is available on those three models only, on the first-party Claude API. Opus 4.6 fast mode was discontinued on June 29, 2026, and Opus 4.7 fast mode was removed on July 24, 2026.
What this means A standard Opus 5.5 request with 10,000 input tokens and 2,000 output tokens costs $0.08. The same request in fast mode costs $0.16, still less than the $0.20 it costs on Opus 5 in fast mode. Use fast mode when response time is the bottleneck, not as a default setting for batch jobs or background processing.
Prompt caching still matters in fast mode because repeated input can be billed at cache-read rates after the fast-mode multiplier. Batch processing is separate: fast mode is not available with the Batch API, so non-urgent work should usually use batch instead of paying the speed premium. Fast mode is an Opus feature; Anthropic does not list it for Fable, Sonnet, or Haiku models.
Prompt Caching
Prompt caching is the single most effective cost feature on the Claude API. When a request reads from the cache, cached tokens are billed at 10% of the standard input rate on most models, a 90% discount. Claude Opus 5.5 and Claude Sonnet 5.5 read at 5% of their input price, $0.20 and $0.10 per million tokens, a 95% discount, and Claude Fable 5.1 and Claude Mythos 5.1 read at 2.5%, $0.25 per million tokens, a 97.5% discount. Claude Haiku 5.5 reads at the standard 10%, which is $0.01 per million tokens on prompts up to 100,000 tokens. Cache writes cost 1.25x the standard input rate for a 5-minute time-to-live, or 2x for a 1-hour time-to-live, on every model.
The break-even math is straightforward. A 5-minute cache pays for itself after a single read. A 1-hour cache pays for itself after two reads. Every subsequent read saves 90% of the normal input cost, 95% on Opus 5.5 and Sonnet 5.5, or 97.5% on Fable 5.1 and Mythos 5.1. The pay-back point is the same on every model; only the saving per read changes. In absolute terms a Sonnet 5.5 cache read now costs $0.10 per million tokens against $0.20 on Opus 5.5, so Sonnet 5.5 costs half as much as Opus 5.5 on every token type, cached or not.

Minimum cacheable prompt length
A prompt has to reach a minimum length before it can be cached at all, and that minimum varies by model. Haiku 5.5, Sonnet 5.5, Opus 5.5, Fable 5.1, Mythos 5.1, Opus 5, Fable 5 and Mythos 5 cache from 512 tokens. Opus 4.8, Sonnet 5, Sonnet 4.6 and Sonnet 4.5 need 1,024. Opus 4.7 and Haiku 3.5 need 2,048. Haiku 4.5, Opus 4.6 and Opus 4.5 need 4,096, the highest floor of any current model, so a prompt of 512 to 4,095 tokens that Haiku 4.5 could not cache can be cached on Haiku 5.5.
Nothing fails loudly when a prompt falls short. The request is processed without caching and no error is returned, so the saving simply does not appear on the bill. Check the response usage fields to confirm: if both cache_creation_input_tokens and cache_read_input_tokens come back as 0, the prompt was not cached. If a prompt falls just short of the threshold, padding the cached section to reach it is usually worth doing.

What to cache
Cache any context that repeats across requests. System prompts, role definitions, tool schemas, reference documents, code examples, and long context that does not change between user messages are all candidates. A chatbot with a 4,000-token system prompt serving 10,000 conversations per day will save roughly $76 daily on Sonnet 5.5 by caching that system prompt, which is about $2,280 per month from a single optimization.
Which TTL to use
Use the 5-minute TTL for interactive workloads where the same context is hit in quick succession, such as chat sessions, multi-turn agents, and iterative coding. Use the 1-hour TTL when cache hits are spread out but predictable within an hour, or when combining cache with batch processing.
Batch Processing
The Message Batches API processes requests asynchronously and returns results within 24 hours, typically under 1 hour, at exactly 50% off both input and output prices. There is no quality difference between batch and real-time responses. The only trade-off is timing.
Batch fits any workload where a few hours of latency is acceptable: document processing pipelines, data enrichment, nightly analytics, offline evaluations, content generation queues, code reviews, and dataset classification. A team processing 500,000 documents per month at Sonnet 5.5 rates saves approximately $500 to $1,500 monthly by routing through batch.
Combining batch with caching
Batch processing and prompt caching compose. A cached read inside a batch request is billed at half the cache-read rate: 5% of the standard input rate on most models, 2.5% on Opus 5.5 and Sonnet 5.5, and 1.25% on Fable 5.1 and Mythos 5.1. On prompt-heavy workloads with repeated context, combined use can reduce input costs by 95% or more. Pair the 1-hour cache TTL with batch rather than the 5-minute TTL, since batch turnaround usually exceeds 5 minutes.
Choosing a Claude Model
Model selection has the largest impact on cost. For high-volume, narrowly scoped work such as classification, extraction, routing and summaries, test Haiku 5.5 first. For other cost-sensitive production workloads, start on Sonnet 5.5 and only move up when benchmarks on your specific workload show a meaningful quality gap. Opus 5.5 covers most complex agentic coding and knowledge work, and Anthropic’s own guidance is to start there for most workloads; it costs less than Opus 5 on every line. Fable 5.1 sits above it at 2.5x the base rate for the most demanding reasoning and long-horizon agents.
When to use Haiku 5.5
Use Haiku 5.5 for short, repetitive, high-volume requests: classification, extraction, routing, moderation, summaries and context compaction, and subagent work under a Sonnet or Opus orchestrator. Anthropic says Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding; on Terminal-Bench 4.0, a long command-line agent test, Haiku 5.5 scored 39.2% against Sonnet 5.5’s 70.6%.
Watch the 100,000-token line. A request with 90,000 input tokens and 1,000 output tokens costs $0.0095; at 110,000 input tokens it costs $0.0575, about six times as much for 22% more input, because every token in the request moves to the higher price. Keep prompts under the line where you can: trim retrieved context, summarize long histories, or split a long document into separate requests. Anthropic has not said whether cached tokens count toward the 100,000, so the calculator above counts them, and so should your budget until Anthropic says otherwise.
Haiku 5.5 uses the newer tokenizer, so the same text counts as roughly 30% more tokens than on Haiku 4.5. Anthropic’s estimate of about 75% lower cost than Haiku 4.5 on average already includes that, and assumes most requests stay under 100,000 tokens, as about 90% of Haiku 4.5 requests did. It is the first Haiku with an effort setting, and the default is medium: low suits chat and simple high-volume requests, high suits knowledge work and longer agent tasks. Check the breaking changes before migrating: manual thinking budgets, non-default temperature, top_p and top_k, and assistant prefill all return errors, computer use needs the new toolset on the Claude API and Google Cloud, and safety classifiers can decline a request with no server-side fallback.
When to use Sonnet 5.5
Use Sonnet 5.5 for most production workloads: everyday coding and bug fixes, tool use, documents, slides and spreadsheets, and high-volume work where speed matters. Its base rates are the same as Sonnet 5, and since October 7, 2026 its cache reads cost half as much, $0.10 against $0.20, so the calculator above shows the same cost for both without caching and a lower one for Sonnet 5.5 on cache reads. Anthropic’s figure of up to 30% less per task comes from Sonnet 5.5 needing fewer tokens, which only your own traffic can show: measure output tokens on a sample after switching, then enter those figures. Against Opus 5.5 it costs half on every line, including cache reads. Anthropic’s own view is that Opus 5.5 remains clearly stronger at complex, open-ended work that needs sustained judgment.
The minimum cacheable prompt is 512 tokens, half of Sonnet 5’s, so prompts between 512 and 1,023 tokens can now be cached. Default effort is high on the Claude API and medium in Claude Code and the Claude apps, and the effort levels are recalibrated, so run a fresh effort sweep rather than carrying Sonnet 5 settings over. Check three breaking changes before migrating: thinking set to disabled returns an error (between_tools is the new lowest setting), forced tool use (tool_choice set to any or tool) returns an error, and on the Claude API and Google Cloud the earlier computer use tool is rejected.
When to use Opus 5.5
Use Opus 5.5 for agentic coding, long-running agents and knowledge work, and as the replacement for Opus 5. Anthropic’s estimate of about 40% lower cost than Opus 5 on typical workloads combines two effects. The rate cut, 20% on input and output and 60% on cache reads, is what the calculator above applies. The rest comes from Opus 5.5 using fewer tokens per task, which no calculator can know in advance: measure output tokens on a sample of your own traffic after switching, then enter those figures.
Default effort is medium, one level below Opus 5’s high, so an integration that never sets effort spends fewer thinking tokens after the switch. Run an effort sweep on your own evals rather than carrying Opus 5 settings over. Check two breaking changes before migrating: thinking cannot be disabled, and forced tool use (tool_choice set to any or tool) returns an error.
When to use Fable 5.1
Use Fable 5.1 when Opus 5.5 at higher effort still falls short: the hardest multi-step reasoning, agents that run unattended for hours, and research or analysis where capability is worth a 2.5x base premium. Two things change the cost picture against Fable 5. Cache reads are $0.25 per million tokens, so a workload that reuses a large prefix on most requests brings its effective input rate down sharply, although output stays at $50 per million tokens against $20 on Opus 5.5; run the cache planner above with your real prefix size and read count to see the blended rate. Effort is the other lever: Anthropic’s launch data shows low and medium effort matching or beating Fable 5 at a fraction of the tokens, so start at medium and raise it only where your evals show the difference. Adaptive thinking is always on and bills at output rates. Plan for refusal handling: a refusal that arrives before any output is billed in some categories (biology, helping develop competing AI models, and reasoning extraction, as of September 2026) and not in others, and a refusal partway through a response bills the tokens already used.
When to stay on Opus 5
Opus 5 costs more than Opus 5.5 on every line, so keep traffic on it only where a workload depends on something Opus 5.5 rejects: forced tool use, or requests that disable thinking. Opus 5.5 also returns the text it writes between tool calls inside thinking blocks rather than text blocks, so an application that streams those progress updates to users needs a display setting change. Pin the version, fix those request shapes, run your evals on Opus 5.5, and migrate once they pass.
When to stay on Sonnet 5
Sonnet 5 costs the same as Sonnet 5.5 on input, output and cache writes and twice as much on cache reads, so keep traffic on it only where a workload depends on something Sonnet 5.5 rejects: thinking switched off with disabled, forced tool use, the computer_20251124 tool on the Claude API or Google Cloud, or the advisor tool with Opus 4.8, Opus 4.7 or Sonnet 5 as the advisor. Pin the version, fix those request shapes, run your evals on Sonnet 5.5, and migrate once they pass.
When to stay on Fable 5
Fable 5 costs the same as Fable 5.1 on base tokens and more on cache reads, and Fable 5.1 scores higher on every benchmark Anthropic published, so the only reason to keep traffic on Fable 5 is a workload you have validated against its specific behavior and have not yet re-evaluated. Pin the version, run your evals on Fable 5.1, and migrate once they pass.
When to stay on Opus 4.8
Opus 4.8 costs the same as Opus 5 and more than Opus 5.5 on every line. Keep it for workloads you have validated against its specific behavior and not yet re-tested; fast mode is available on Opus 5.5 at the same 2x multiplier, so speed is not a reason to stay.
When to stay on Haiku 4.5
Haiku 4.5 costs more than Haiku 5.5 on every line at any prompt length: $1 / $5 against $0.10 / $0.50, or $0.50 / $2.50 above 100,000 tokens. Keep traffic on it only where a workload depends on something Haiku 5.5 rejects, such as assistant prefill, temperature or top_p settings, manual thinking budgets, or Priority Tier capacity. Its retirement is not sooner than October 15, 2026, so plan the move now: pin the version, fix those request shapes, run your evals on Haiku 5.5, and migrate once they pass.
Output tokens cost more than input
Output is billed at five times the input rate across every current model. A request with 5,000 input tokens and 5,000 output tokens on Sonnet 5.5 costs $0.01 for the input and $0.05 for the output, so the output represents 83% of the bill. Keep responses concise. Request specific output formats such as JSON or bullet points when the task allows. Set length constraints when detailed analysis is not required. Thinking tokens are billed as output tokens, and on models where thinking is always on, such as Opus 5.5 and Fable 5.1, effort is the main control over that volume. Monitor thinking token usage on long-running agents.
Claude Haiku 5.5 Benchmarks
Scores from Anthropic’s Haiku 5.5 announcement on October 7, 2026, against Haiku 4.5, OpenAI’s GPT-6 Luna and, for reference, Sonnet 5.5. OSWorld 2.1 is scored on its offline subset, and a dash means no published score. Haiku 5.5 is far ahead of Haiku 4.5 and ahead of GPT-6 Luna on every row where they have a score, but well behind Sonnet 5.5 on Terminal-Bench 4.0, which is why Anthropic positions it for narrowly scoped work rather than long agentic coding.
Three Sonnet 5.5 figures here differ from the Sonnet 5.5 table below: GDPval-AA v2.1 (1840 here, 1844 on September 28), AA-Briefcase v1.1 (1824 here, 1811 then) and computer use (83.9% on the OSWorld 2.1 offline subset here, 80.1% on the partial set then). Anthropic has not explained the knowledge-work changes; the September table noted that those runs used a pre-release deployment with a since-fixed bug. Each table shows the figures as that announcement published them.

Claude Sonnet 5.5 Benchmarks
Scores from Anthropic’s Sonnet 5.5 announcement on September 28, 2026, against Sonnet 5, Opus 5.5 and OpenAI’s GPT-6 Sol. On FrontierCode, Sonnet 5.5 scored 52.1% at xhigh effort and 46.2% at max, because at max it more often ran a multi-agent code review, which in two examined cases caused a timeout or edits beyond the task’s scope. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment with a structured-output bug that has since been fixed; Anthropic expects any effect to be small and to understate Sonnet 5.5. A dash means no published score. Anthropic also reports cost per task: on Terminal-Bench 4.0, Sonnet 5.5 at medium effort beats Sonnet 5’s best score for less than a tenth of the cost per task.
The computer-use benchmark is labelled OSWorld 2.1 here and OSWorld 2.0 in the Opus 5.5 table below, with the same 81.8% for Opus 5.5. Chartography here is scored without tools; the Opus 5.5 table reports it with tools, which is why the Opus 5.5 figures differ.

Claude Opus 5.5 Benchmarks
Scores from Anthropic’s Opus 5.5 announcement on September 22, 2026. Opus 5.5 ran with adaptive thinking at max effort, except Terminal-Bench 4.0, which is reported at xhigh. Production safeguards were on; where they intervened, cybersecurity tasks were completed by Opus 4.8 and biology tasks by Opus 5, which likely lowers the Opus 5.5 figures. GPT-6 Astra scores are as reported by OpenAI, and a dash means no published score. Anthropic also reports cost per task: at its default medium effort, Opus 5.5 scored 54.6% on FrontierCode, above GPT-6 Astra’s best of 53.3%, for about a fifth of the cost per task.
Two Fable 5.1 figures here differ from Anthropic’s September 1 announcement in the next section: OSWorld 2.0 (80.7% here, 77.9% then) and Humanity’s Last Exam with tools (65.6% here, 65.0% then). Anthropic has not explained the change, so each table shows the figures as that announcement published them.

Claude Fable 5.1 Benchmarks
Scores from Anthropic’s Fable 5.1 announcement on September 1, 2026, with Fable 5 and Opus 5 re-run under the same conditions. Fable 5.1 was evaluated with its production safeguards on; where a safeguard intervened, the task scored zero, so these are floor figures for the model. Mythos 5.1 is the same model without those interventions and scored 60.9% on Terminal-Bench 4.0 against Fable 5.1’s 55.8%. OSWorld 2.0 uses the August 2026 task release, which is why no competitor score is shown there.

Claude Fable 5 Benchmarks
Benchmark comparison from Anthropic’s Fable 5 announcement. Anthropic reports Fable 5 and Mythos 5 scores jointly; on starred benchmarks Fable 5 can land closer to Opus 4.8 because its safety fallbacks route some queries away.

Claude Sonnet 5 Benchmarks
Benchmark comparison from Anthropic’s Claude Sonnet 5 launch on June 30, 2026. Sonnet 5 sits clearly above Sonnet 4.6 across every headline benchmark and approaches Opus 4.8, edging ahead of Opus 4.8 on GDPval-AA v2 knowledge work. Opus 4.8 is shown as the reference ceiling.

Additional API Charges
Base token costs are not the only line item on an API bill. Several features carry their own pricing.
Web search
$10 per 1,000 searches for the search tool itself. The input and output tokens produced by processing the search results are billed separately at the model’s standard rate.
Code execution
Code execution is free when the request also includes the web search or web fetch tool. Used on its own it is billed by execution time: each organization receives 1,550 free container-hours per month, and usage beyond that runs at $0.05 per hour per container, with a 5-minute minimum per execution.
Managed Agents sessions
$0.08 per session-hour of active runtime. Idle time while the session waits for the next message does not count. Standard token rates still apply on top, and prompt caching, fast mode and US residency pricing apply inside sessions; the batch discount does not.
US-only data residency
For workloads that must run inside the US, specifying US-only inference via the inference_geo parameter adds a 1.1x multiplier across every token category, covering input, output, cache writes, and cache reads. Applies to Claude 4.6 and later, so Haiku 5.5, Sonnet 5.5, Opus 5.5, Fable 5.1, Mythos 5.1, Fable 5, Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5 and Sonnet 4.6, on the first-party Claude API and Claude Platform on AWS. Earlier models, including Haiku 4.5, Sonnet 4.5 and Opus 4.5, return an error if the parameter is sent. Third-party platforms have their own regional pricing.
Tool use overhead
Tool use adds a small system prompt automatically, and the size depends on the model and the tool choice: 286 tokens on Opus 5.5 and Sonnet 5.5, which accept only the auto and none tool choices, 286 to 406 on Haiku 5.5, 286 to 474 on Opus 5, Opus 4.8 and Sonnet 5, 496 to 589 on the 4.5 and 4.6 models, and 675 to 804 on Opus 4.7. The text editor tool adds around 700 tokens, and bash adds 325 on Opus 5, 4.8 and 4.7 or 244 on Opus 4.6 and earlier; Anthropic has not published a bash figure for Opus 5.5, Sonnet 5.5 or Haiku 5.5. The overhead is per-request, so it matters most on high-volume, short-message workloads.
1M Context Window
Sonnet 5.5, Opus 5.5, Fable 5.1, Mythos 5.1, Fable 5, Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, and Sonnet 4.6 include the full 1 million token context window at standard flat pricing. A 900,000-token request is billed at the same per-token rate as a 9,000-token request. This is a change from earlier generations, where long-context requests above 200,000 tokens were billed at higher rates.
Haiku 5.5 is the exception among current models. It also has a 1M-token window, but a prompt over 100,000 tokens is billed at five times the rate on every token in the request, input, output and cache alike. The calculator above switches to the higher price when the input tokens per request pass 100,000.
Prompt caching and batch discounts apply at standard rates across the full context window. For applications that analyze large codebases, long legal documents, or entire books in a single pass, the flat pricing keeps spend predictable.
Planning Token Counts
Claude’s tokenizer produces roughly 1.33 tokens per word for English prose, or about 4 characters per token. A planning rule of thumb: 1,000 words is approximately 1,330 tokens. Technical documentation runs slightly higher at around 1,400 tokens per 1,000 words. Source code runs higher still at around 1,500 tokens per 1,000 words, because punctuation and syntax tokenize less efficiently than prose.
Those ratios describe the older tokenizer. Anthropic states that Claude 4.7 and later models, which include Haiku 5.5, Sonnet 5.5, Opus 5.5, Opus 5, Fable 5.1 and Sonnet 5, produce approximately 30% more tokens for the same text, with the exact increase depending on the content. The words and characters options in the calculator use the older ratio, so on those models budget from real token counts where you can.
For accurate counts before deployment, the Anthropic API includes a token counting endpoint that reports exact token usage for any input. The estimator above applies the 4-characters-per-token heuristic, which is close enough for budgeting but not exact. Re-run token counts when migrating between models because identical text can produce different billable token totals across model versions.
Best Cost Control Practices

A few operational practices keep API spend predictable as usage grows.
Track token usage per request from the start. The API response includes a usage object reporting exact input and output token counts. Monitor per-feature and per-user spend rather than only the monthly total, since the total alone will not identify which feature or customer is responsible for a spike.
Cache anything that repeats. System prompts, tool schemas, documentation embedded in the prompt, and few-shot examples are all candidates. Any context that appears in more than one request should be cached.
Route non-urgent requests through batch. Anything that does not need a sub-second response, such as nightly jobs, content generation pipelines, and evaluation runs, belongs on the Batch API at half price.
Set spending limits in the Anthropic Console. Organization-level and workspace-level limits prevent runaway spend from a bug or a misconfigured agent. Alerts at 50%, 75%, and 90% of the monthly budget give time to react before hitting the hard cap.
Benchmark smaller models on your actual task. Teams often default to Opus when Sonnet or even Haiku would suffice. Running a representative sample of real workload through each tier and measuring the quality gap empirically is usually worth the evaluation time, because the cost difference is large.
Pricing verified from Anthropic’s pricing documentation.
See also the detailed pricing documentation and model overview.