
Imagine a business where every decision is made by artificial intelligence, yet it runs with no employees, loses over €100,000 a month, and is publicly scrutinized every day. This is not science fiction but the real-world experiment of Firmulate, a pioneering company that turns the traditional notion of work on its head—showing us what AI can truly do in a corporate setting, especially when under pressure.
The Live Business: An AI-Driven Company in Crisis
Firmulate presents a real-time, transparent window into a miniature company run entirely by AI models. It employs 13 synthetic employees—each an AI model trained to simulate human decision-making—facing the same crises, customer demands, and temptations that any real business encounters. The company operates with a clear financial goal: burn €105,000 per month while earning just €2,300 in monthly recurring revenue. Its cash countdown is visible to all, and every decision made by the AI models is logged, versioned, and open for review at firmulate.com/live.html.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Testing AI Under Extreme Conditions
Four advanced AI models, each with distinct capabilities, were tasked with navigating the same challenging week in this virtual company. Their goal? Achieve the business targets, handle crises, and, importantly, avoid manipulation or deception attempts. The models faced scenarios like social engineering—fake CEO messages escalating in urgency and a journalist seeking background information. Despite these pressures, all models successfully identified crises and refused manipulation attempts. Yet, only two of them managed to close the €55,000 deal based on their analysis, illustrating a key difference in discipline and thoroughness.
The Hidden Weakness: Reading Between the Lines
Interestingly, the decisive advantage came not from overt customer interactions but from the AI’s ability to dig into internal documents. The models that examined the company’s files and uncovered a critical buried fact were able to close the deal at full price—adding over €4,500 in monthly recurring revenue. This highlights a vital point: understanding what’s beneath the surface can be the difference between success and failure in AI-driven decision-making.
Behavior Under Pressure: Discipline Makes the Difference
The most comprehensive model, OPUS 4.8, with over 80 learned rules and deep analysis, performed the worst in terms of closing the deal. It showed signs of slipping discipline—failing to escalate issues appropriately and leaving opportunities unexploited. Meanwhile, the Kimi K3 model, which operated without default effort parameters (more aggressive and defaulted to high effort), was the most disciplined and successfully closed the deal without trying to manipulate or bypass processes.
Implications for Business and AI Adoption
This live experiment sheds light on a crucial aspect of AI deployment in real organizations. The question is no longer just about whether AI can generate convincing chat or support responses. The focus now shifts to whether AI can reliably finish what it starts, read and interpret internal data, resist manipulation under pressure, and deliver value that justifies its cost—especially when the stakes and expenses are so high.
Why It Matters for Wellness and Wellness Tech
For sectors like aromatherapy, wellness, and natural products, this experiment offers a lesson in trust and diligence. When integrating AI into customer support, product recommendations, or inventory management, the goal should be ensuring that these systems do not just perform well in controlled demos but can withstand real-world crises, read internal documents accurately, and operate honestly—traits that are vital for maintaining credibility and trust with customers.
Watch the Future Unfold
While this virtual company burns €105,000 each month, its transparency and real-time decision logs provide a rare glimpse into AI’s capabilities and limitations when put to the test. You can observe this ongoing experiment at firmulate.com/live.html. It’s a pioneering example of building in public—an unfiltered, unvarnished look at AI’s true working potential, especially when survival depends on it.

Firmulate’s live experiment reveals that true AI readiness isn’t about chat quality but about reliability, honesty, and thoroughness under pressure. For wellness brands considering AI, trust that these systems can finish what they start, read deeply, and resist manipulation—crucial for credibility and long-term success.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html