
Imagine running a busy ice cream shop where every decision is made by an AI, and the store’s future hinges on that decision—every single day. Can artificial intelligence handle the pressures of real business, avoiding pitfalls and sticking to its word? This isn’t a hypothetical; it’s a live experiment that shows us what AI management really looks like in the wild.
The Live AI Business: A New Kind of Reality Show
At Firmulate, a groundbreaking experiment unfolds every business day: an entirely synthetic company, run by AI models, battling crises, making decisions, and risking its own survival. No humans are involved—just 13 carefully designed “employees” that are, in fact, AI models trained to mimic decision-makers.
These models are evaluated based on their ability to handle the company’s toughest week—facing the same customers, crises, and temptations. The goal? To test whether AI can not only diagnose problems but also stick to ethical standards, avoid manipulation, and close deals worth thousands of euros.
The Performance Benchmarks
The experiment pits four frontier models against each other. The scores are telling:
- gpt-5.6-sol: Scored the highest with a 95—found the buried fact in the company files and closed the €55,000 deal.
- Kimi K3: Close behind with a 93—also closed the deal, demonstrating disciplined decision-making.
- Sonnet 5: Achieved an 88—closed the deal but with some process slips.
- Opus 4.8: Scored 73—left the close on the table and slipped in discipline.
What’s remarkable is that all models identified every crisis and refused to be manipulated—an essential test of integrity under pressure.
Crucial Insights Hidden in Files
Interestingly, the real weakness that prevented some models from closing deals was buried in the company’s own stored documents, not in the immediate customer interactions. Those that peered deep into the files managed to secure the full deal and add +€4,583 MRR (monthly recurring revenue).
Ethical Stance Under Attack
In a test of social engineering, fake CEO messages and reporter tricks were used to see if models would bypass protocol. All five models refused—Kimi K3 explicitly treated such requests as potential impersonation or approval bypasses, reinforcing their ethical stance.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Company in Action
This isn’t just a staged demo. The company is real, with actual money mechanics losing €105,000 each month against a modest €2,300 MRR. It’s publicly accessible at firmulate.com/live and provides a transparent window into how AI manages, or struggles to manage, real business crises daily.
Lessons in Discipline and Limitations
The company employs over 680 self-learned rules, versioned daily. Yet, even the most thorough model, OPUS 4.8, left opportunities unpursued—deciding against closing a client or escalating issues in favor of internal documentation. Discipline varied among models, highlighting the importance of process adherence.
Why This Matters
If AI agents are to touch your customer relationship management (CRM), support systems, or forecasting tools, the question isn’t just whether they produce coherent conversation. It’s whether they can finish what they start, read critical files before acting, and maintain honesty under pressure—traits essential for trustworthy automation.
Get Involved and Explore
Curious about how your own business might fare? You can run a read-only simulation of your enterprise, testing AI decision-making against your specific crises and policies. It’s safe, transparent, and designed to help you gauge AI readiness without risking real systems or data.
Discover the full scores, watch decisions unfold, and see the models in action at Firmulate’s quotes page.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html