
By Thomas Cohen, founder of Maestro
Quality, measured: six days of duels against the reference method
We promised that one day we would measure more than restraint. We have: over 180 jury verdicts, twelve method versions, an initial defeat, and a lesson about measurement itself.
In July, the article about our ×10 ended with a promise: if we ever measure beyond injected context, there will be another article with the same level of detail. Here it is. Between August 5 and 10, we had the method directing Maestro's agents judged against BMAD-METHOD v6, the reference open method, on what ultimately matters: the quality of the delivered product.
The protocol
Nine realistic projects, identical on both sides: expense claims with VAT, shared expenses for housemates, a time clock, room bookings, invoicing, inventory, a library, habit tracking and file organisation. Each method starts from the same brief and proceeds to a verified product. Three judges grade each submission with the same rubrics and anonymised entries. They compile, run and probe the code instead of settling for the documents.
Internally, our method is called Partition. The name carries Maestro's thesis: short documents to perform, rather than a manual to recite.
We started by losing
The first duel, in early August, ended at 82.1 versus 90.8: a clear defeat. The verdicts showed where: tricky cases never exercised, retroactivity, duplicates, cents lost to rounding, documents contradicting the code, and missing business rules. Every subsequent version came from these verdicts. Twelve versions in six days, each placing a rule exactly where the previous one had failed: exercise money-related cases from the specification stage, prove each repair in real conditions, and inject domain expertise, multiple VAT rates, rounding remainders, payroll, only when the project touches it.
What the figures say
At the end of the campaign, using the same rubrics and judges: 93.3 out of 100 for Partition, 83.5 for the reference method. On cost, across the nine end-to-end projects, Partition crossed the finish line nine times out of nine for $115.65; the reference exhausted its budget three times out of nine before delivering, for $238.66. Its design is rich and its documents often more extensive than ours; the gap widens in execution, where the delivered code remains thin and tests are rarely run. Our original restraint has not changed: the method did not grow by a single byte over seven versions, and the ×10 context ratio still holds.
The lesson that makes this article worthwhile
The latest version returned an average of 87.0, six points below its predecessor. We nearly concluded there had been a regression. Instead, we organised paired duels: the same judge, two submissions for the same project side by side, anonymised and randomly ordered. Verdict: the new version wins six duels out of eight. Probes showed why those averages were misleading. Several attractive earlier scores rewarded documents sitting atop code that had never been probed: a time clock rated 92.5 did not compile and wiped its database at every launch; a 91.6 silently overwrote files. Recent submissions deliver much more executable code; the judges run it, and their strictness has shifted with the material they receive.
Since then, we have adopted a rule: averages can only be compared within the same campaign, and between two versions, only paired duels count. Mapped back to the original scale through those duels, the current version is worth around 96 out of 100.
What the figure still does not tell you
These scores come from our own protocol, with AI juries, not an independent laboratory. A single run per project leaves noise of five to nine points; our conclusions rest on structural changes and duels, never decimal places. Measured defects also remain on the register: the protocol develops only the first batch of each project, and wiring a module to its screen caused trouble in three consecutive campaigns. These two issues are the next work item.
The exercise cost around $1,280 over six days, for more than 180 jury verdicts and twelve engine versions. Every rule in every version answers a measured defect. Today's figure reads as follows: higher quality on the same rubrics, half the cost, and the finish line crossed every time.