Journal
Journal: behind the scenes
On this topic, 17 articles from the journal, from practical guides to behind the scenes of the product.

Application development is conducted like an orchestra
A score written before playing, sections with clear remits, a conductor holding the baton without playing an instrument: the metaphor structuring Maestro, embraced all the way to its logo.
Read the article
Why agency prices have not followed the fall in AI costs
What cost $60 per million tokens in late 2021 now costs between $0.06 and $0.40 at equivalent performance, a reduction of 150 to 1,000 times. Meanwhile, a developer's average daily rate in France remains around €629. Someone is keeping that margin.
Read the article
Self-service skills: the supply chain nobody is watching
The github/spec-kit repository shows 130,054 stars a year after launch. The AI agent skills ecosystem is growing just as fast, and installing a skill on demand is becoming the new curl piped into a terminal.
Read the article
What our beta testers broke, second pass: the silence that blocked everything
After a first real user test, 21 written comments revealed a defect invisible internally: an entire step type was handled nowhere in the engine. The chain froze silently, waiting for an answer that never came.
Read the article
What does an application built by agents really cost? Our bills, line by line
Unlimited AI plans are dying, and budgets are exploding at businesses without caps. We publish our real costs: $105.64 for a complete campaign of 5 judged tasks, from brief to verified product.
Read the article
“Looks good to me”: why human approval is AI’s real safeguard
Article 14 of the European AI Act requires ‘effective’ human oversight of high-risk AI. At Maestro, nothing is agreed without ‘Looks good to me’: not a regulatory constraint added later, but the product's central mechanism from day one.
Read the article
Which drifted, the system or the judge? We measured our jury’s strictness
An evaluation of 21 AI judges shows rankings shifting by up to 14 places depending on the benchmark. Our average score dropped six points between campaigns: we had to prove whether our engine or our jury had changed.
Read the article
Therapists and care professionals: confidentiality as the number-one requirement
On November 6, 2025, Doctolib was fined €4.6 million for abuse of a dominant position. In June 2026, health data was being used to train AI models. For a therapist, the question is no longer which software, but where their notes live.
Read the article
Our benchmarks publish our defeats: what SWE-bench can no longer tell you
32.7% of successful SWE-bench fixes contain solution leakage, according to a recent position paper. As the reference leaderboard collapses, we publish our own duels, including BOTH victories and defeats.
Read the article
Can an AI agent team deliver your software? What we measured
Nine projects taken from idea to verified product, three judges per submission, costs published to the cent. What measurement says about an agent team's real capabilities, and what it does not.
Read the article
Why we left BMAD: an autopsy in numbers
Maestro started with BMAD-METHOD. We left it for our own engine, Partition, and measured the replacement campaign after campaign: an average of 89.5 versus 83.5 across nine tasks, with eight duels won out of nine.
Read the article
Our harshest tester is an agent: five temperaments, from rushed to destructive
Since July, every new engine version has faced an agent playing a user with five temperaments, from rushed to destructive, capped at $15: five real bugs caught on day one.
Read the article
First week of beta: what broke, what we fixed
Our first tester never saw the second screen. An honest account of the first days: three real defects, what they taught us, and what remains open.
Read the article
Quality, measured: six days of duels against the reference method
We promised that one day we would measure more than restraint. We have: over 180 jury verdicts, twelve method versions, an initial defeat, and a lesson about measurement itself.
Read the article
What our ×10 measures, and what it does not
We publish one figure: around 10 times less methodological context injected into agents. Here is the full measurement, including its limits.
Read the article
The deliberate opposite of a terminal
Why Maestro looks like a concert hall rather than a black screen: the setting is part of the product.
Read the article
Why nine agents have first names
Margaux, Victor, Salomé, Maurice, Constance, Marcel, Amélie, Félix and Léa. Names that make the work easy to follow.
Read the article