
By Thomas Cohen, founder of Maestro
Subscriptions, credits and tokens: understanding AI usage
Why do your assistant and provider show different figures? Learn to read the units, check calculations and connect the numbers to a working session.
Your assistant displays thousands of tokens, a usage percentage or a few cents. The provider shows a different amount. Before deciding that a counter is wrong, check what each figure measures: a response, a session, an account, a period or an invoice. This guide helps you reconcile the information without mixing units.
We will follow a fictional example: Nadia is building an order-tracking tool for her business. She uses an AI assistant through Maestro and wants to understand one working session. The quantities, amounts and calculations below are invented to explain the method. They are neither customer bills nor prices charged by Maestro or a provider.
Subscription, API, credits and tokens: four separate concepts
A subscription is the plan you have purchased. It may provide access to tools, models and a certain amount of usage under its terms. An API is a way for software to communicate with a service. It is not another word for a subscription, nor a guarantee that your usual subscription pays for that usage.
A credit is a unit defined by a commercial offer. Its value and rules depend on the service. One tool’s credit cannot be compared directly with another’s. A token is a unit of model processing: it does not represent one completed action in your application.
OpenAI keeps ChatGPT and API billing separate. ChatGPT usage credits are not API credits. Check the relevant plan and sign-in method before interpreting a balance.
Identify the account and the authoritative record
Start by identifying the service responsible for the usage. For Nadia, “I paid for AI” is not precise enough. She records the assistant, connected account, organization if applicable, and sign-in method. She can then open the relevant record instead of searching through all her subscriptions.
Claude Code’s documentation explains that session cost is an estimate and authoritative billing comes from the provider. On a subscription, the session cost figure does not necessarily represent an additional amount payable.
Cursor’s usage tracking separates included usage and on-demand charges. Mistral’s usage report allows you to examine details such as the period, model and workspace involved.
- Record the account and organization that actually ran the task.
- Identify the service and plan to which that usage belongs.
- Open the corresponding usage history, then the invoice when investigating a billed amount.
- Note the period, currency and whether the figure is estimated or final.
Provider pages change. Keep the consultation date and plan name with your record, without copying an API key or password. This gives you something verifiable if the display or terms change the following month.
Understand input, output and reused material
A token is not always a word. The division depends on factors including text, language and model. Some reasoning units are not visible in the answer, as explained in OpenAI’s token documentation.
When reading a record, distinguish what the model receives from what it produces. In our example, Nadia writes a short request, but the assistant may also need to process the order rules, relevant files and results of its checks. Nadia’s message length therefore does not describe the whole session.
- Input: material sent to the model so it can work.
- Output: material produced by the model, under the categories counted by the provider.
- Cache: certain reused material, with processing and pricing that depend on the service.
Anthropic’s caching documentation distinguishes cache writes and reads, each with its own conditions. Reused material should therefore not automatically be treated as free input or added without checking which total already includes it.
Work through a complete calculation with fictional prices
Take an invented teaching example: €2 per million ordinary input tokens, €0.20 per million cache-read tokens and €8 per million output tokens. These figures exist only to explain the calculation. They do not correspond to a recommended offer and exclude taxes and other services.
A fictional request uses three distinct categories: 50,000 ordinary input tokens, 100,000 cache-read tokens and 5,000 output tokens. For each category, divide the quantity by one million, then multiply it by the corresponding rate. Finally, add the three amounts.
- Ordinary input: 50,000 ÷ 1,000,000 × €2 = €0.10.
- Cache reads: 100,000 ÷ 1,000,000 × €0.20 = €0.02.
- Output: 5,000 ÷ 1,000,000 × €8 = €0.04.
- Total for this fictional request: €0.16.
One user request may involve several exchanges
If our invented session contains six exchanges exactly like this one, the total reaches 6 × €0.16, or €0.96. In real work, quantities can change with every exchange. Multiplying the first cost by the number of visible messages therefore does not automatically reconstruct the final amount.
The calculation changes if a line concerns cache writes, another model, a paid tool or a different unit. Also check whether rates are quoted per thousand or per million tokens. Using the wrong unit produces the wrong result even when the multiplication itself is correct.
Keep the record’s currency throughout reconciliation. If you later want to convert the result, record the exchange rate and date separately. Mixing a dollar price with a euro bank payment from the start makes the difference harder to explain.
Why the tool and provider may show different figures
Two figures can both be correct while covering different scopes. One relates to the open session; the other includes all activity on the account. One screen tracks the last few minutes, another the billing month. Align those scopes before looking for an error.
Maestro’s tracking distinguishes subscription mode from API mode and uses information reported by the assistant. Available details therefore depend on that assistant. A missing or old measurement does not prove that ongoing work consumes nothing. Also consult the relevant provider’s usage record.
- Period: the same start time, end time and time zone.
- Scope: the session, project, user or entire organization.
- Type: estimated usage, recorded cost, consumed credits or an invoice.
- Unit: currency, tokens, a percentage or credits from the same plan.
Imagine Nadia sees €2.40 for a session and €3.10 on the account over the same period. The fictional €0.70 difference might come from other work, but that is only a hypothesis. She looks for matching entries before reaching a conclusion. A plausible explanation is not yet a reconciled record.
Finally, check how recent the data is: a total may update after an operation finishes. Record the time of each screenshot and compare again once data labeled as provisional has settled. Do not rerun a task merely to see whether the counter moves.
On a subscription, do not guess what a percentage costs
A usage percentage is not automatically a percentage of your subscription price. If a gauge shows 30%, that does not establish an additional charge equal to 30% of the monthly fee. You need to know which allowance it measures and the period over which it renews.
In Nadia’s fictional case, a session may stay entirely within included usage. It consumes part of an allowance without necessarily creating a new charge. If the plan permits paid overage, check the relevant settings and record instead of inferring overage from the activity displayed.
Likewise, do not treat an equivalent calculated at API rates as a subscription invoice. Such an equivalent can help compare approximate quantities within a defined protocol. It does not replace the plan’s terms or establish what will be charged.
Reconcile a session in six steps
For your first reconciliation, choose a short session with a known objective. Nadia selects a fix to the calculation for cancelled orders. She avoids starting with several weeks, assistants and devices at once. A small scope makes it easier to identify entries without inventing connections.
- Write down the session’s goal, start time and end time.
- Record the assistant, known model, account and billing method.
- Save the tool’s counter and its last measurement time, if shown.
- Open the provider’s record for the same dates and correct account.
- Identify other account activity, cost categories and possible rounding.
- Classify the remaining difference as explained, partly explained or requiring investigation.
Suppose the €0.70 from the earlier example turns out to belong to a second session in another tool. Nadia can explain the difference without adjusting the first counter. If no entry matches, she leaves the difference unresolved. She does not invent a “technical margin” to force the totals to agree.
When asking for help, prepare the plan name, period, amounts, known model and screenshots with sensitive information removed. A request identifier may help when available. Do not send keys, customer data or confidential content simply to illustrate a difference between counters.
Include retries and attempts that produce no usable result
A failed task may still have consumed model work. Usage does not measure result quality. Keep abandoned attempts in your record when analyzing a session, while distinguishing them from validated outcomes.
In another fictional example, Nadia makes six attempts to obtain three verified fixes. Each attempt costs €0.16 here, for a total of €0.96. The average cost per verified fix is therefore €0.32. Counting only the three successful attempts would hide half the work consumed.
That number is still incomplete on its own. Correcting a label and repairing an important calculation are different outcomes. Also record the type of task, successful checks and human time required. You can then compare similar sessions without turning averages into promises.
Keep a useful record without accounting for every token
One row per session is often enough to begin: date, objective, assistant, account, available measurement, result and an explanation of any difference. Keep token-level detail only when it helps explain an anomaly or compare options under the same conditions.
For Nadia, a useful entry might read: “Cancelled orders; rule corrected; three checks passed; API cost reconciled; twenty minutes of review.” If no per-task amount is available, she says so. The record preserves the connection between consumption and actual progress.
Keep this record separate from estimating an application’s overall budget. Hosting, maintenance, outside services and team time answer different questions. Our guide to application development costs helps examine those items without confusing them with one AI session.
Reduce retries without stripping useful detail from the request
To improve usage, first look for unnecessary work: an undecided rule, contradictory changes or a request repeated without understanding the failure. A shorter prompt is not automatically better. Removing an essential constraint can cause more rework than supplying it would have required.
Define an expected outcome, provide the relevant material and check one step before requesting the next. For selecting an engine by task, see our guide to AI assistants and models in Maestro. A comparison matters only if you also measure quality and the work left for you.
For an example of measurements presented with their scope, read the record of our agent trials. It describes a specific campaign. Its amounts do not set your project’s price or replace your own records.
Start with one session you can explain from beginning to end. Maestro is free during its invitation-only Mac beta; AI usage remains separate. You can explore Maestro and request an invitation, then keep this habit: identify the unit, account and period before judging a figure.
Read the complete guide: build an application without coding