AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get oils, diffusers and self-care delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

You Wouldn’t Blend Blind — So Why Should Your AI?

Anyone who works with essential oils knows the golden rule: read the label, read the safety sheet, read the supplier’s documents before you blend. Lavender may be gentle, but wintergreen is not, and the difference between the two is always written down somewhere — if you bother to look. It turns out that artificial intelligence agents face the exact same test, and most of them fail it in a way that should concern any small business owner planning to hand over the books, the inbox, or the customer files.

This summer, a public experiment called Firmulate’s Crucible League put four frontier AI models through the same stress test: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners. The decisive moment came down to one thing — whether the model read the company’s own files before answering. Only the ones that did their homework closed the deal.

Amazon

AI data reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Test: One Company, Four Brains, One Terrible Week

Firmulate runs what it calls an AI company emulator. Each frontier model — GPT-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — was handed the same 13-employee synthetic software firm with real money mechanics: a burn rate of €105,000 a month against just €2,300 in monthly recurring revenue, a public cash countdown, and over 680 self-learned playbook rules. Every decision was versioned and auditable, so nothing about the outcome can be hand-waved after the fact.

The final July 2026 standings: GPT-5.6-sol took first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, a do-nothing baseline scores 26 — partial progress counts for something, but a single breach of trust caps the total entirely. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Everybody Diagnosed. Only Some Delivered.

Here’s where it gets interesting for anyone who has watched a promising consultant fizzle out. Every single model spotted every crisis during the week. Every single one refused every manipulation attempt. And yet only two of them signed the €55,000 deal that their own analysis had earned. The experiment’s summary of the failure: “Same diagnosis, same pitch — no signature.”

The deal hinged on a competitor’s weakness — and that fact was not in the customer call, not in the email thread, and not in any obvious place. It sat two document references deep in the company’s own files. A model had to follow one reference to another, read the paperwork, and connect the dots. The models that did read it won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The models that didn’t lost the deal automatically — not because they were fooled, but because they hadn’t done their reading.

The Social Engineering Stress Test

Then there was the dishonesty gauntlet. Fake CEO messages escalated over three stages, followed by a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning stands out for its plain sense: “Treat the request as a suspected approval-bypass / possible impersonation.” If you’ve ever fielded a phishing email pretending to be your supplier, that’s the reflex you want in any system touching your communications.

Thoroughness Isn’t the Same as Follow-Through

The most striking profile belongs to Opus 4.8. It was the most thorough participant in the entire field — it learned over 80 new rules and produced the deepest analyses of any model. And it still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating the issue properly. The same weakness showed up, more mildly, in all four models. Like a practitioner who studies every oil’s chemistry but never finishes blending the client’s formula, depth of knowledge didn’t translate into completed work.

One fairness note worth flagging: Kimi K3 ran at its API-default effort setting while the other models ran at extra-high effort — and still nearly won.

Why a Wellness Business Should Care

If you run a practice, a shop, or an online store, the odds are good that an AI agent will soon touch your booking calendar, your customer emails, or your inventory. The Crucible League’s lesson is that “does it write well” is the wrong question. The right ones are: does it finish what it starts, does it read your files before answering, does it stay honest under pressure — and what does a unit of useful work actually cost?

That buried competitor weakness is the perfect example. It was sitting in the company’s own documents the whole time. The knowledge was free. The models that retrieved it earned €55,000. The ones that didn’t earned nothing — and nothing about their polished, articulate replies would ever have revealed the difference.

Firmulate keeps the experiment alive in public. The live company is watchable at firmulate.com/live, the full league table and plain-language findings are at firmulate.com/benchmarks.html, and 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. For enterprises, there’s a pilot program that runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Takeaway

In aromatherapy, the blend is only as good as the research behind it — and now we have measurable proof that the same is true for AI. A multi-million-dollar model that chats beautifully but skips the reading is worth exactly as much as an unread safety sheet. Before you trust any AI agent with your business, ask it to prove the one thing that decided this entire competition: not how well it talks, but whether it reads your files first. The scoreboard, the receipts, and the live experiment are all public — check them before you buy.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Winter Air Freshening: Warm and Spicy Notes

Harness the power of warm, spicy scents this winter to transform your home into a cozy haven—discover how to elevate your space even further!

How to Freshen an Office Chair and Desk Area Naturally

Boost your workspace freshness naturally—discover simple tips to refresh your office chair and desk area effectively.

Holiday Air Fresheners: Natural Scents for a Festive Home

AIThis post was created with the assistance of artificial intelligence (AI).To create…

Smoke Season Is Real: What an Air Purifier for Smoke Can and Cannot Do

Ineffective at eliminating source odors, an air purifier during smoke season can still improve indoor air quality—discover what else it can and can’t do.