
By Thomas Cohen, founder of Maestro
Can an AI agent team deliver your software? What we measured
Nine projects taken from idea to verified product, three judges per submission, costs published to the cent. What measurement says about an agent team's real capabilities, and what it does not.
Yes, provided you direct it. Across nine projects taken from idea to verified product, Maestro's agent team scored 93.3 out of 100 before a shared jury, versus 83.5 for the reference open method, crossing the finish line nine times out of nine for $115.65 in tokens.
First, what we mean by ‘agent team’
An AI agent is a model assigned a narrow role: scoping a need, writing a specification, building a stage or checking a result. An agent team sequences those roles with human approval points between them. The awkward question: does this arrangement produce working software or attractive demonstrations? We wanted a measured answer rather than an opinion.
The protocol, without jargon
Nine realistic projects of the kind an SME commissions: expense claims with VAT, shared expenses, a time clock, room bookings, invoicing and inventory. Two methods start from the same brief: ours and BMAD-METHOD v6, the field's reference open method. Each submission proceeds to the final product, then three judges grade it against the same rubrics, with submissions anonymised. Instructions to the judges: compile, run and probe the code instead of merely reading the documents.
The results, including defeat
The first duel ended in defeat: 82.1 versus 90.8. The verdicts identified cases never exercised, retroactivity, duplicates, cents lost to rounding, and documents contradicting code. We corrected the method following every verdict, twelve versions in six days. At campaign end: 93.3 versus 83.5, and the finish line crossed nine times out of nine, while the reference exhausted its budget three times along the way, for $238.66.
The anecdote that makes measurement worthwhile
One time-clock submission received 92.5 out of 100 on the strength of its documents. When a judge ran its code, it did not compile and wiped its database on every startup. Remember this detail when choosing a tool: AI-delivered software is judged by running it. Our judges run everything; demand the same of what is delivered to you.
What the measurement does not say
These scores come from our own protocol, with AI juries, not an independent laboratory. One run per project leaves noise of five to nine points between executions; our conclusions rest on clear gaps and paired duels, never decimal places. And nine management projects prove nothing about a high-load critical system: for that, a specialist agency remains the right contact.
The answer to the title's question
An agent team delivers real, verified management software at a cost within an SME's budget, subject to one condition with no measured exception in our work: someone directs it. You approve the requirements, decide the rules and reject what is wrong. All nine completed projects share that point. The protocol and its limits are described in detail in our benchmark articles.