Two small handwritten cards on a desk, one covered in ticks and the other in crosses, with a miniature brass trophy placed exactly between them, on neither card

August 23, 2026 · 6 min read

By Thomas Cohen, founder of Maestro

Our benchmarks publish our defeats: what SWE-bench can no longer tell you

32.7% of successful SWE-bench fixes contain solution leakage, according to a recent position paper. As the reference leaderboard collapses, we publish our own duels, including BOTH victories and defeats.

A recent position paper identifies solution leakage in 32.7% of successful fixes on SWE-bench, the reference leaderboard for coding agents (arxiv.org/pdf/2606.17799). As the industry's reference collapses under its own biases, we publish our own duels on tasks invented for each campaign, victories and defeats included.

What broke trust

SWE-bench is not the only concern. A 2025 Stack Overflow survey measures 46% distrust in AI tools' accuracy among developers against 33% trust, while 84% use them anyway (survey.stackoverflow.co/2025/ai). And a Cohere/Stanford/MIT/Allen AI study showed that Meta tested 27 private variants of its models on the LMArena leaderboard before Llama 4, publishing only the best (arxiv.org/abs/2504.20879). The problem is no longer an isolated bug; it is an established practice, explaining why a figure displayed alone without a verifiable protocol behind it no longer convinces anyone serious.

What we do differently

Every campaign starts from new tasks, never published before the run, judged by three judges on both sides with the same rubrics and anonymised submissions. We began in early August by losing: our first duel ended at 82.1 versus 90.8 for the reference method. That result remained public. Every subsequent engine version corrected exactly what that jury had criticised, until the trend reversed: the complete details, with the initial defeat, are in quality measured over six days of duels.

We have felt the temptation to cherry-pick

Publishing the best run and keeping the others quiet is a real temptation, not a vice we attribute to others from a distance. Our safeguard: the same judges assess both sides of the duel in the same randomly drawn presentation order, and each campaign's dollar costs are published beside the scores. A score without its cost, and without the question ‘did it actually finish?’, says nothing.

What we gain by showing a defeat

A documented defeat is a transferable lesson; a victory alone is not. In a recent campaign, an apparently excellent score, 92.5 out of 100, actually rewarded a project that did not compile and wiped its own database at every launch. Only by digging deeper, probing executed code rather than reading the document describing it, did the defect emerge. Publishing this kind of result, with an explanation of what was poorly measured and why, is worth infinitely more than a round figure without context, however flattering it might be to us in the short term.

What you can check yourself

Detailed reports for every campaign, task by task, with full judges' verdicts and dollar costs, exist internally and inform every article we publish about our method. We do not ask to be taken at our word; we ask to be checked against evidence, exactly as we do for what our ×10 restraint figure measures. A benchmark its reader cannot challenge is not a benchmark: it is an advertisement dressed in numbers.

Back to the journal

Take the baton.

Leave your email to try Maestro in the first waves.

The beta opens in waves. People on the list try it first, and Maestro stays free throughout the beta.

The beta is currently available on macOS 13 or later. Your answer helps us plan other versions.

Your email is only used to let you know when access opens. Nothing else, we promise.