A kitchen table in the evening with an open account book, calculator and still-warm cup

September 3, 2026 · 6 min read

By Thomas Cohen, founder of Maestro

Claude Fable 5.1 pricing: the reduction is in the cache, not the headline rate

Anthropic introduced Claude Fable 5.1 on September 1. The per-token rate is unchanged; rereading context falls by 75%. In practice, long building sessions that keep the same project in view for hours become the cheapest way to work.

Claude Fable 5.1 pricing takes three lines at Anthropic, which introduced it on September 1, 2026: $10 per million input tokens, $50 for output, and $0.25 to reread what the model has already read, 75% less than the previous version.

$10 and $50per million tokens, input and output (Anthropic, September 1, 2026)
$0.25rereading previously read context, 75% less than Fable 5
around 25%Anthropic's estimated saving on a typical workload

The cache in one sentence

When an assistant works on your project, it reloads the same file at every exchange: your requirements, files already written and yesterday's decisions. The cache is the part the provider sets aside instead of charging again at the input rate, then rereads at a reduced price. At Anthropic, this rereading falls to $0.25 per million tokens, four times less than before, while input and output rates remain unchanged. A ten-minute conversation will not notice. A session lasting a morning on the same project pays this line with every response.

What that means on a bill

Anthropic estimates savings of around 25% on a typical workload and up to around 45% for uses where an agent performs many stages on its own (official Claude Fable page, September 1, 2026). These are its own estimates, not an independent measurement, and depend on a figure the page does not provide: the share of your spending spent rereading the same context. A tool starting from scratch with every question gains almost nothing. A tool building software for three hours while keeping the same requirements in view gains a lot.

Long sessions become the rewarded use

This pricing changes what is expensive. A short session, where you ask one question and leave, pays the full rate on everything it sends. A long session, where an agent team writes a specification, breaks it into stages, then builds and checks each one, repeatedly sends the same documents: that is the share falling to $0.25. A method working for a long time on a stable file therefore benefits more than a sequence of isolated questions at equal token volume. We publish our real costs campaign after campaign, and this line weighs heavily in the total.

What the reduction does not solve

Reuse only works if the beginning of the submitted context stays identical between exchanges. A tool reordering files, inserting a different instruction at the start each turn, or opening a new session for every question pays the input rate again without saying so. You will see no difference on screen: responses arrive, quality remains, only the bill changes. And the saving concerns reading, not production: an agent writing lots of code still pays $50 per million output tokens.

Anthropic had released Claude Opus 5 on July 24, 2026 at $5 for input and $25 for output, unchanged from the previous generation. Two models, two rate cards, and in both cases the per-token rate cannot predict spending: the number of exchanges needed to finish the task matters more than unit price. The same mechanism explains why agency prices have not followed falling model costs, and why an AI budget is managed through caps rather than rate-card comparisons.

The counter that counted incorrectly

Your tool must still count this line. Ours did not: until August 2026, Maestro's token counter ignored cached rereading and displayed a total matching no bill. We rebuilt it to count rereading at its price. A tester could not see it, the total looked plausible, and that makes this kind of defect slow to find: you must retrieve the provider's statement and place it beside your own screen, line by line, to discover the two are not talking about the same thing.

Check your own usage

Before comparing two tools on token price, ask each what it charges to reread your project and check whether its counter separates that line from the rest. If it shows only an overall total, run the same task twice consecutively on the same project: the second session should cost less than the first. If the amount does not change, you are paying full price on text the model has already read, and the right contact for that question is the tool's vendor, not the model provider.

Back to the journal

Take the baton.

Leave your email to try Maestro in the first waves.

The beta opens in waves. People on the list try it first, and Maestro stays free throughout the beta.

The beta is currently available on macOS 13 or later. Your answer helps us plan other versions.

Your email is only used to let you know when access opens. Nothing else, we promise.