
Picture an ice cream shop on its busiest Saturday: a freezer fails, a supplier calls, a customer complains, and someone claiming to be the owner asks for an exception. A polished chatbot can sound helpful through all of it. The harder question is whether an AI entrusted with real work will make the right calls when the pressure lands.
Get baking supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate’s experiment puts that question to the test. Its public live company is watchable, and its enterprise pilot offers a way to wargame crisis scenarios against a business using a read-only export.
Same bad week, different models
In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable.
The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s stated rule is that partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Spotting trouble is only half the job
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. They reached the same diagnosis and made the same pitch; most still left the signature on the table.
The deal hinged on a detail hidden in the company’s own files, two document references deep. It was not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding is a reminder that an AI can recognize a problem and still miss the context that makes a useful response possible.
The pressure tests included fake CEO messages escalating over three stages, plus a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work still needs a close
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and discipline slipped: it attempted to write into a locked department instead of escalating. The same weakness appeared, in weaker form, across all four.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The standings are the reported outcome of this particular experiment.
A live company, then a company-specific pilot
The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. Its experiment is watchable at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess the model.
For an enterprise, the next step is less like watching a company from the outside and more like rehearsing with its own material. The pilot uses a read-only export to create a digital twin, then runs crisis scenarios against the business and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

Practice before the pressure is real
The experiment suggests that recognizing a crisis and refusing a bad request are not enough: models also need to find relevant evidence, follow company discipline, and finish work they have already justified. A company-specific wargame can make those gaps visible before an AI handles real operations.
To discuss an enterprise pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
