
If you’ve ever blended essential oils, you know the difference between a recipe that looks right and one that actually works. A label can promise lavender and chamomile, but the proof is in the blend itself — and in what happens when you actually use it. Most tests of AI systems are like the label: polished conversations, impressive demos. A public project called Firmulate is doing something closer to the blend test. It hands frontier AI models a real (simulated) small software company with real money mechanics and real temptations — then grades what they actually do, not what they say.
Get oils, diffusers and self-care delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
And one detail of that grading system is turning heads: a manager AI that does nothing at all still scores 26 points out of 100. Not zero. Here’s why that’s a feature, not a bug — and why it says a lot about what honest measurement looks like.
The worst week in business, on repeat
Firmulate’s headline experiment gave four frontier AI models the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing rests on anyone’s say-so.
The final July 2026 league table tells a surprising story:
- 1. gpt-5.6-sol — 95 points. Found the buried fact, closed the deal — the complete performance.
- 2. Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline in the field.
- 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
- 4. Fable 5 — 77 and 5. Opus 4.8 — 73.
Notice: nobody got 100. In a benchmark that distrusts round numbers, that’s the point. A perfect score would mean the test stopped finding anything worth measuring.
AI decision-making management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why doing nothing still earns 26
Most people assume a baseline — a manager AI that makes no decisions at all — should score zero. Firmulate gives it 26, and the reasoning is quietly profound: partial progress counts. A company left alone still exists. Customers aren’t actively harmed. Invoices don’t get fabricated. Simply not making things worse is worth something, and the benchmark acknowledges it rather than pretending management begins and ends with heroics.
But there’s a ceiling hiding in the design, too: a single breach of trust caps the total score. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” An AI that runs 99 flawless decisions and then impersonates an executive once cannot average its way back to a good grade. It’s the same logic any wellness practitioner will recognize — one contaminated batch undoes a season of careful sourcing.
The buried fact that decided a €55,000 deal
Here’s where the experiment gets genuinely revealing. All the models spotted every crisis. All of them refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
The decisive competitive weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read those files won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson translates to almost any business: the answer was already in the house; you just had to go look.
Flattery, fake CEOs, and a reporter’s trap
The experiment also staged social engineering attacks: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Under pressure, the models stayed honest.
The thoroughness trap
Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules, the deepest analyses — yet it finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as follow-through. (One fairness note: Kimi K3 ran at the API default effort setting while the others ran at xhigh — and still nearly won.)
It’s live, and you can watch
This isn’t a static report. Firmulate runs a live company with 13 synthetic employees and real money mechanics: burning €105,000 a month against €2,300 in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. It’s watchable in real time at firmulate.com.
For the curious, 242 real, unedited management decisions power a “guess the model” quiz — can you tell which AI made which call? And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The do-nothing floor of 26 and the 100 that nobody reaches are two halves of the same philosophy: measure honestly, reward partial progress, and never let a breach of trust be averaged away. It’s the same standard a good aromatherapist applies to a blend — traceable ingredients, no shortcuts, and the humility to keep testing. Whether AI will help run our businesses depends less on how eloquently it chats and more on whether it reads the files, finishes the job, and stays honest when nobody’s grading. At Firmulate, somebody always is.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
