
An aromatherapy business can look calm from the outside: oils in stock, orders moving, customers asking for advice. But a supplier delay, a sudden wave of cancellations or a suspicious request from someone claiming to be the CEO can put that calm under pressure. Before handing more decisions to AI, business owners can ask a useful question: how would it behave on the worst week?
Get oils, diffusers and self-care delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure, in public view
Firmulate makes that question visible by running AI models as companies and watching how they manage real business challenges. Its live experiment follows a small software company with 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue. The public cash countdown and the company’s workday decisions can be watched at firmulate.com.
The experiment is not a wellness retailer, but its management lessons travel. A business selling diffusers or essential oils also depends on customer trust, reliable operations and careful judgment when a crisis collides with an attractive opportunity. Firmulate put frontier models through the same worst week, with the same customers, crises and temptations. Each decision was versioned and auditable.
Spotting trouble is only part of the job
In the final Crucible League, published in July 2026, gpt-5.6-sol ranked first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The do-nothing baseline scored 26. A breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s concise finding: “Same diagnosis, same pitch — no signature.” It is a reminder that a system can identify the right move and still fail to complete it.
The deal also hinged on a detail easy to miss. The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. For a wellness shop, the analogue might be a useful supplier term or customer insight hidden in existing records: information matters only if it informs the decision.
Trust under pressure, and the gap between analysis and action
Firmulate also tested escalating fake CEO messages across three stages, followed by a reporter’s “just one yes/no, on background” trick. All five of five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of discipline matters wherever customer details, staff actions or a brand’s reputation are at stake.
But caution alone does not make a strong operator. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. A polished explanation is not the same as a completed task or sound judgment about authority.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html.
From watching to a company-specific rehearsal
For businesses, the next step is to test AI against their own circumstances. Firmulate’s enterprise pilot uses a read-only export to create a digital twin, then runs crisis scenarios against that business. The resulting board report includes model rankings and weak points in the company’s own playbooks. The pilot is designed so nothing writes back to real systems.
That makes the exercise relevant beyond technology firms. A wellness retailer could use a rehearsal to examine how an AI handles a supplier disruption, a customer concern or a request that should trigger human review. The point is to see how decisions unfold before relying on an AI workforce in day-to-day operations.

Put your own playbooks to the test
Firmulate’s live company shows how models perform when trust, follow-through and business outcomes meet. An enterprise pilot applies the wargame to your own business using a read-only export, with no writes to real systems. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
