
By Thomas Cohen, founder of Maestro
Claude Fable 5.1 pricing: the reduction is in the cache, not the headline rate
Anthropic introduced Claude Fable 5.1 on September 1. The per-token rate is unchanged; rereading context falls by 75%. In practice, long building sessions that keep the same project in view for hours become the cheapest way to work.
Claude Fable 5.1 pricing takes three lines at Anthropic, which introduced it on September 1, 2026: $10 per million input tokens, $50 for output, and $0.25 to reread what the model has already read, 75% less than the previous version.
The cache in one sentence
When an assistant works on your project, it reloads the same file at every exchange: your requirements, files already written and yesterday's decisions. The cache is the part the provider sets aside instead of charging again at the input rate, then rereads at a reduced price. At Anthropic, this rereading falls to $0.25 per million tokens, four times less than before, while input and output rates remain unchanged. A ten-minute conversation will not notice. A session lasting a morning on the same project pays this line with every response.
What that means on a bill
Anthropic estimates savings of around 25% on a typical workload and up to around 45% for uses where an agent performs many stages on its own (official Claude Fable page, September 1, 2026). These are its own estimates, not an independent measurement, and depend on a figure the page does not provide: the share of your spending spent rereading the same context. A tool starting from scratch with every question gains almost nothing. A tool building software for three hours while keeping the same requirements in view gains a lot.
Long sessions become the rewarded use
This pricing changes what is expensive. A short session, where you ask one question and leave, pays the full rate on everything it sends. A long session, where an agent team writes a specification, breaks it into stages, then builds and checks each one, repeatedly sends the same documents: that is the share falling to $0.25. A method working for a long time on a stable file therefore benefits more than a sequence of isolated questions at equal token volume. We publish our real costs campaign after campaign, and this line weighs heavily in the total.
What the reduction does not solve
Reuse only works if the beginning of the submitted context stays identical between exchanges. A tool reordering files, inserting a different instruction at the start each turn, or opening a new session for every question pays the input rate again without saying so. You will see no difference on screen: responses arrive, quality remains, only the bill changes. And the saving concerns reading, not production: an agent writing lots of code still pays $50 per million output tokens.
Anthropic had released Claude Opus 5 on July 24, 2026 at $5 for input and $25 for output, unchanged from the previous generation. Two models, two rate cards, and in both cases the per-token rate cannot predict spending: the number of exchanges needed to finish the task matters more than unit price. The same mechanism explains why agency prices have not followed falling model costs, and why an AI budget is managed through caps rather than rate-card comparisons.
The counter that counted incorrectly
Your tool must still count this line. Ours did not: until August 2026, Maestro's token counter ignored cached rereading and displayed a total matching no bill. We rebuilt it to count rereading at its price. A tester could not see it, the total looked plausible, and that makes this kind of defect slow to find: you must retrieve the provider's statement and place it beside your own screen, line by line, to discover the two are not talking about the same thing.
Check your own usage
Before comparing two tools on token price, ask each what it charges to reread your project and check whether its counter separates that line from the rest. If it shows only an overall total, run the same task twice consecutively on the same project: the second session should cost less than the first. If the amount does not change, you are paying full price on text the model has already read, and the right contact for that question is the tool's vendor, not the model provider.