
By Thomas Cohen, founder of Maestro
Token prices fall, bills rise: what AI agents cost and how to count it
Gartner forecasts a more than fivefold increase in inference costs per agent workflow by 2028, while token prices collapse. The two trends go together: an agent that plans uses 5 to 30 times more tokens than a simple chatbot. Here is what to look for on your own bill.
The cost of AI agents on your bill rises while the unit price falls. Gartner forecasts that inference costs per agent workflow will more than quintuple by 2028, while token prices are expected to fall 95% by 2030 (Gartner, August 17, 2026, reported by Computerworld the same day).
Why both trends run together
A chatbot receives a question and answers. An agent reads a document, chooses an action, executes it, rereads the result, corrects itself, and starts again. Each loop sends the entire context again. Analysts Will Sommer and Sabine Zimmerhansl, quoted by Computerworld, put numbers to the gap: an agent capable of reasoning costs up to 150 times more than a basic assistant for a single task, consumes 5 to 30 times more tokens, and its provider charges 8 to 10 times the simple-task rate for a planning token. The unit price really is falling; the number of units per completed job is growing faster.
What the user sees at the other end
A user of an application generation platform did the sums and published them on Reddit (r/lovable, September 2026): 1,364 credits consumed in thirty days, 2,389 over three months, meaning 57% of the quarter's usage came in the final month. They measured twenty exchanges, 281.4 credits in total, ranging from 8.9 to 21.1 credits each, averaging 14.07. Twenty interactions had therefore eaten 70% of their monthly allocation of 400 credits. Their question was less about the amount than about the lack of visibility: nothing in the interface explained what was consuming it.
The industry is beginning to address the problem through limits rather than explanations. On August 31, 2026, Vercel launched per-user spending budgets for its AI gateway, rejecting new requests once the limit is reached, with alerts at 50%, 75%, and 100% of budget and a dedicated administrative role. A spending cap has become a marketable feature, which says a lot about how often unpleasant surprises occur.
Counting accurately, including what gets reused
Our own token counter was wrong until August 2026. It ignored the cache, the part of the context the provider retains between exchanges and charges at a different rate. Users therefore saw a number below their bill, the worst mistake a counter can make: it reassured them. We rebuilt it in August 2026 to include the cache, and since then the screen has matched what the provider charges. A counter that underestimates is worth less than no counter, because it stops people asking the question. The lesson extends beyond our case: before trusting a number displayed in an interface, ask the publisher what it counts and what it ignores. Cache, failed attempts, and automatic reviews are the three items counters most often forget.
The three limits to require before launching an agent workflow
The first is an estimate before launch, with a multiplier bounding expenditure: in our product, a project stops at three times its estimate. The second is a monthly budget across all projects, independent of what happens within each one. The third is a limit on attempts: when a step fails verification, two repairs are attempted, then the decision returns to you instead of letting the loop continue. We publish our actual costs, line by line, including the repair budget before it was capped.
Read your bill tonight
Open your provider's console and look for three things: your cost over the last seven days, the share of input tokens versus output, and whether there is a cap and which address receives its alerts. If the last seven days cost more than the previous month divided by four, something changed in your usage and nobody told you. Look for changes in your habits before blaming the provider: a longer conversation, a larger document attached to every exchange, or a task restarted three times can explain a doubling. Our analysis of the price reduction hidden in the cache explains why the input line deserves attention first, and annual maintenance remains the expense these consoles will never show.
Read the complete guide: build an application without coding