A woman directs an AI agent team from her desk in the evening, with Maestro's tracking board on screen

August 23, 2026 · 6 min read

By Thomas Cohen, founder of Maestro

Can an AI agent team deliver your software? What we measured

Nine projects taken from idea to verified product, three judges per submission, costs published to the cent. What measurement says about an agent team's real capabilities, and what it does not.

Yes, provided you direct it. Across nine projects taken from idea to verified product, Maestro's agent team scored 93.3 out of 100 before a shared jury, versus 83.5 for the reference open method, crossing the finish line nine times out of nine for $115.65 in tokens.

93.3 out of 100the directed agent team, before a shared jury
83.5 out of 100the reference open method, same rubrics
$115.65the total cost of the nine completed projects

First, what we mean by ‘agent team’

An AI agent is a model assigned a narrow role: scoping a need, writing a specification, building a stage or checking a result. An agent team sequences those roles with human approval points between them. The awkward question: does this arrangement produce working software or attractive demonstrations? We wanted a measured answer rather than an opinion.

The protocol, without jargon

Nine realistic projects of the kind an SME commissions: expense claims with VAT, shared expenses, a time clock, room bookings, invoicing and inventory. Two methods start from the same brief: ours and BMAD-METHOD v6, the field's reference open method. Each submission proceeds to the final product, then three judges grade it against the same rubrics, with submissions anonymised. Instructions to the judges: compile, run and probe the code instead of merely reading the documents.

The results, including defeat

The first duel ended in defeat: 82.1 versus 90.8. The verdicts identified cases never exercised, retroactivity, duplicates, cents lost to rounding, and documents contradicting code. We corrected the method following every verdict, twelve versions in six days. At campaign end: 93.3 versus 83.5, and the finish line crossed nine times out of nine, while the reference exhausted its budget three times along the way, for $238.66.

The anecdote that makes measurement worthwhile

One time-clock submission received 92.5 out of 100 on the strength of its documents. When a judge ran its code, it did not compile and wiped its database on every startup. Remember this detail when choosing a tool: AI-delivered software is judged by running it. Our judges run everything; demand the same of what is delivered to you.

What the measurement does not say

These scores come from our own protocol, with AI juries, not an independent laboratory. One run per project leaves noise of five to nine points between executions; our conclusions rest on clear gaps and paired duels, never decimal places. And nine management projects prove nothing about a high-load critical system: for that, a specialist agency remains the right contact.

The answer to the title's question

An agent team delivers real, verified management software at a cost within an SME's budget, subject to one condition with no measured exception in our work: someone directs it. You approve the requirements, decide the rules and reject what is wrong. All nine completed projects share that point. The protocol and its limits are described in detail in our benchmark articles.

Back to the journal

Take the baton.

Leave your email to try Maestro in the first waves.

The beta opens in waves. People on the list try it first, and Maestro stays free throughout the beta.

The beta is currently available on macOS 13 or later. Your answer helps us plan other versions.

Your email is only used to let you know when access opens. Nothing else, we promise.