
By Thomas Cohen, founder of Maestro
Our benchmarks publish our defeats: what SWE-bench can no longer tell you
32.7% of successful SWE-bench fixes contain solution leakage, according to a recent position paper. As the reference leaderboard collapses, we publish our own duels, including BOTH victories and defeats.
A recent position paper identifies solution leakage in 32.7% of successful fixes on SWE-bench, the reference leaderboard for coding agents (arxiv.org/pdf/2606.17799). As the industry's reference collapses under its own biases, we publish our own duels on tasks invented for each campaign, victories and defeats included.
What broke trust
SWE-bench is not the only concern. A 2025 Stack Overflow survey measures 46% distrust in AI tools' accuracy among developers against 33% trust, while 84% use them anyway (survey.stackoverflow.co/2025/ai). And a Cohere/Stanford/MIT/Allen AI study showed that Meta tested 27 private variants of its models on the LMArena leaderboard before Llama 4, publishing only the best (arxiv.org/abs/2504.20879). The problem is no longer an isolated bug; it is an established practice, explaining why a figure displayed alone without a verifiable protocol behind it no longer convinces anyone serious.
What we do differently
Every campaign starts from new tasks, never published before the run, judged by three judges on both sides with the same rubrics and anonymised submissions. We began in early August by losing: our first duel ended at 82.1 versus 90.8 for the reference method. That result remained public. Every subsequent engine version corrected exactly what that jury had criticised, until the trend reversed: the complete details, with the initial defeat, are in quality measured over six days of duels.
We have felt the temptation to cherry-pick
Publishing the best run and keeping the others quiet is a real temptation, not a vice we attribute to others from a distance. Our safeguard: the same judges assess both sides of the duel in the same randomly drawn presentation order, and each campaign's dollar costs are published beside the scores. A score without its cost, and without the question ‘did it actually finish?’, says nothing.
What we gain by showing a defeat
A documented defeat is a transferable lesson; a victory alone is not. In a recent campaign, an apparently excellent score, 92.5 out of 100, actually rewarded a project that did not compile and wiped its own database at every launch. Only by digging deeper, probing executed code rather than reading the document describing it, did the defect emerge. Publishing this kind of result, with an explanation of what was poorly measured and why, is worth infinitely more than a round figure without context, however flattering it might be to us in the short term.
What you can check yourself
Detailed reports for every campaign, task by task, with full judges' verdicts and dollar costs, exist internally and inform every article we publish about our method. We do not ask to be taken at our word; we ask to be checked against evidence, exactly as we do for what our ×10 restraint figure measures. A benchmark its reader cannot challenge is not a benchmark: it is an advertisement dressed in numbers.