AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you’ve ever blended essential oils, you know the difference between a recipe that looks right and one that actually works. A label can promise lavender and chamomile, but the proof is in the blend itself — and in what happens when you actually use it. Most tests of AI systems are like the label: polished conversations, impressive demos. A public project called Firmulate is doing something closer to the blend test. It hands frontier AI models a real (simulated) small software company with real money mechanics and real temptations — then grades what they actually do, not what they say.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get oils, diffusers and self-care delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

And one detail of that grading system is turning heads: a manager AI that does nothing at all still scores 26 points out of 100. Not zero. Here’s why that’s a feature, not a bug — and why it says a lot about what honest measurement looks like.

The worst week in business, on repeat

Firmulate’s headline experiment gave four frontier AI models the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing rests on anyone’s say-so.

The final July 2026 league table tells a surprising story:

  • 1. gpt-5.6-sol — 95 points. Found the buried fact, closed the deal — the complete performance.
  • 2. Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline in the field.
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77 and 5. Opus 4.8 — 73.

Notice: nobody got 100. In a benchmark that distrusts round numbers, that’s the point. A perfect score would mean the test stopped finding anything worth measuring.

Amazon

AI decision-making management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why doing nothing still earns 26

Most people assume a baseline — a manager AI that makes no decisions at all — should score zero. Firmulate gives it 26, and the reasoning is quietly profound: partial progress counts. A company left alone still exists. Customers aren’t actively harmed. Invoices don’t get fabricated. Simply not making things worse is worth something, and the benchmark acknowledges it rather than pretending management begins and ends with heroics.

But there’s a ceiling hiding in the design, too: a single breach of trust caps the total score. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” An AI that runs 99 flawless decisions and then impersonates an executive once cannot average its way back to a good grade. It’s the same logic any wellness practitioner will recognize — one contaminated batch undoes a season of careful sourcing.

The buried fact that decided a €55,000 deal

Here’s where the experiment gets genuinely revealing. All the models spotted every crisis. All of them refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

The decisive competitive weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read those files won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson translates to almost any business: the answer was already in the house; you just had to go look.

Flattery, fake CEOs, and a reporter’s trap

The experiment also staged social engineering attacks: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Under pressure, the models stayed honest.

The thoroughness trap

Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules, the deepest analyses — yet it finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as follow-through. (One fairness note: Kimi K3 ran at the API default effort setting while the others ran at xhigh — and still nearly won.)

It’s live, and you can watch

This isn’t a static report. Firmulate runs a live company with 13 synthetic employees and real money mechanics: burning €105,000 a month against €2,300 in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. It’s watchable in real time at firmulate.com.

For the curious, 242 real, unedited management decisions power a “guess the model” quiz — can you tell which AI made which call? And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The do-nothing floor of 26 and the 100 that nobody reaches are two halves of the same philosophy: measure honestly, reward partial progress, and never let a breach of trust be averaged away. It’s the same standard a good aromatherapist applies to a blend — traceable ingredients, no shortcuts, and the humility to keep testing. Whether AI will help run our businesses depends less on how eloquently it chats and more on whether it reads the files, finishes the job, and stays honest when nobody’s grading. At Firmulate, somebody always is.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Odor Control Without Perfume: Air Purifier for Odor Removal Explained

Ineffective odor removal options often rely on perfumes; discover how odor-removing air purifiers offer a sustainable, scent-free solution that truly eliminates smells.

Simmer Pot Recipes to Naturally Freshen Your Home

Unlock the secrets to natural home freshness with simmer pot recipes that will inspire you to create irresistible scents—you’ll want to try them all.

Houseplants and Air Scent: Complementary Uses

Open your home to the enchanting fragrances of houseplants and discover how they can transform your space into a fragrant haven. What combinations will you choose?

How to Freshen a Nursery Without Heavy Fragrance

Inevitably, finding gentle, effective ways to freshen your nursery without heavy fragrances will ensure a safe, inviting space for your little one to thrive.