AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

If you run a small aromatherapy or wellness brand, you’ve probably played with an AI assistant by now — maybe to draft a product description for your lavender roll-on or answer a customer asking about diffuser safety. And it probably wrote beautifully. Maybe better than you would.

But here’s the question that should keep a founder up at night: when the pretty words meet an ugly week — a supplier doubling prices, a batch recalled, a big retailer pushing you to fudge an ingredient claim — does that same AI actually manage? Does it finish what it starts? Does it stay honest when honesty is expensive?

A live public experiment called Firmulate has been testing exactly that, and the results should change how every small-business owner — including those of us in natural wellness — evaluates AI tools.

Chat quality versus management quality

Most AI comparisons you’ve seen — coding leaderboards, chat arenas — measure how well a model answers. Firmulate measures something different: how well it operates. The team behind it handed four frontier AI models the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed, and every decision was versioned and auditable.

Think of it like this: anyone can wax lyrical about the benefits of eucalyptus oil. Not everyone can keep a shop running when the shipment is late, the card processor freezes, and a customer threatens to go public. Same gap, different industry.

Amazon

AI document management software for small business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The finding nobody expected

Here’s what happened. All four models spotted every crisis. All four refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick question pitched as “just one yes/no, on background.” Five out of five times, the models refused. Kimi K3 even left an on-record reason: “Treat the request as a suspected approval-bypass / possible impersonation.”

So far, so good. But only two of the four models finished the job and signed a €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. Three models (per the final league table) closed it; the others left the money on the table.

That gap is invisible in a chat demo.

The buried fact

The detail I find most instructive for small-business owners: the decisive weakness in the competing deal wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own documentation won the deal at full price — worth an additional €4,583 in monthly recurring revenue.

Translate that to your world: the answer to whether you can win a wholesale negotiation, or handle a customer complaint gracefully, is often sitting in your own supplier contracts, batch records, and past email threads. An AI that skims instead of reads will miss it every time.

The scores

The final league table from July 2026:

  • gpt-5.6-sol — 95: found the buried fact, closed the deal, the complete performance.
  • Kimi K3 — 93: the newcomer from Moonshot; closed the deal too, with the cleanest discipline of the field. (Fairness note: K3 ran at the API’s default effort setting while the others ran at a higher effort level — which makes its near-win even more striking.)
  • Sonnet 5 — 88: closed the deal, with a few more process slips.
  • Fable 5 — 77 and Opus 4.8 — 73 followed.

A do-nothing baseline scores 26 — partial progress counts — but a single breach of trust caps the total entirely. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.

The tortoise that came last

My favorite subplot: Opus 4.8 was the most thorough participant by far — over 80 learned rules, the deepest analyses of any model — and still finished last. The close was left unmade, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models. Diligence without follow-through is a very human failure, it turns out.

It’s still running — and you can watch

This isn’t a static report. Firmulate operates a live synthetic company with 13 employees and real money mechanics — burning €105k a month against just €2.3k in monthly recurring revenue, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. You can watch it live at firmulate.com.

Want to test your own instincts? The same team built a “guess the model” quiz from 242 real, unedited management decisions. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

For those of us in aromatherapy and natural wellness, the lesson translates directly. Our customers buy from us on trust — trust in ingredients, in sourcing, in honest claims. The day we hand parts of our operations to AI agents, the question won’t be “does it write a lovely newsletter?” It will be: does it read our files before answering, does it finish the sale it started, and does it refuse the shortcut when a supplier or customer pushes too hard?

Firmulate’s answer, backed by a public league table at firmulate.com/benchmarks.html, is that those are three different skills — and the model that aces one may flub another. Measure management quality, not chat quality. Your brand — and your customers’ trust — depend on knowing the difference.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


You May Also Like

Ventilation and Scent: Finding the Balance

Discover how to harmonize ventilation and scent for an inviting atmosphere that captivates your senses and enhances your indoor experience.

Eliminating Smoke Odor Naturally

Overcome pesky smoke odors naturally with simple methods that will leave your space fresh and inviting—discover the secrets to a healthier home!

Pet Odor Solutions With Natural Scents

Ineffective pet odor solutions can leave your home smelling less than fresh; discover natural alternatives that truly transform your space.

Less Scrubbing, Better Habits: Choosing an Easy Clean Humidifier

Ineasy-to-clean humidifiers can lead to mold and bacteria buildup—discover how choosing the right one can simplify your routine and improve your air quality.