A hand sorts receipts one by one beneath a brass lamp, beside an open account book covered in handwritten figures

August 28, 2026 · 6 min read

By Thomas Cohen, founder of Maestro

What does an application built by agents really cost? Our bills, line by line

Unlimited AI plans are dying, and budgets are exploding at businesses without caps. We publish our real costs: $105.64 for a complete campaign of 5 judged tasks, from brief to verified product.

Unlimited AI plans are dying, and businesses are switching to cheaper models to stay within budget (tomshardware.com, article on rising AI costs). In contrast, we publish our real costs campaign after campaign: $105.64 for a complete campaign of 5 judged tasks, including specification, development and verification, without truncation.

$105.64a complete campaign of 5 tasks, zero truncation
around $30repair budget measured, then capped, per work package
around $37two real crash-test fleets
$15spending cap for the automated test harness

What pricing articles do not show

Content about AI development costs discusses individual developer subscriptions, a few dozen dollars per month per person. Nobody shows what a complete feature costs, from approved brief to verified code, including the jury. Our published figures fill that gap campaign after campaign, using the same documents we drew on to measure our method against BMAD.

A campaign in detail

A complete measurement campaign, five tasks run in parallel through to verified product and assessed by several judges on both sides, cost $105.64, without any of the five exhausting its budget before completion. That is the price of a full measurement pass, not a single project: divided by five, the order of magnitude per task falls below $25.

The repair budget, measured then capped

When a story fails verification, a targeted repair is triggered, followed by another check. This mechanism has a measured cost, around $30 per work package before capping, and is governed by a strict rule: at most two repairs per work package; beyond that, the decision returns to the human instead of looping. A budget growing without limit is not a technical detail but a product risk, and the same logic bounds an agent's attempts before it stops simply repeating the same thing, hoping every time for a different, better result.

Crash testing and the harness cap

Having an agent with a destructive temperament confront the engine across two real test fleets cost around $37 in total. And every code push triggers an automatic harness pass, capped at $15 per session, so no cost drift goes unnoticed before even reaching production.

The product mechanics that protect your budget

These caps are not confined to our testing behind the scenes: they exist in the product itself. An estimate multiplied by three serves as a cap before development begins, a monthly budget bounds all your projects and a queue prevents multiple work packages from consuming resources in parallel without control. The principle is the same everywhere: a visible figure before spending, never after.

What this changes for you

These amounts do not predict what your project will cost: every product has its own complexity, and more ambitious work naturally consumes more tokens and time. But they provide a verifiable order of magnitude, published campaign after campaign, rather than a sweeping price promise or an agency quote that never details what it truly covers. A figure accompanied by its calculation method can be discussed and checked; a figure standing alone merely asks to be taken on trust. During the free, invitation-only beta, additional costs amount to a few euros in artificial intelligence tokens.

Back to the journal

Take the baton.

Leave your email to try Maestro in the first waves.

The beta opens in waves. People on the list try it first, and Maestro stays free throughout the beta.

The beta is currently available on macOS 13 or later. Your answer helps us plan other versions.

Your email is only used to let you know when access opens. Nothing else, we promise.