AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Picture an ice cream shop on its busiest Saturday: a freezer fails, a supplier calls, a customer complains, and someone claiming to be the owner asks for an exception. A polished chatbot can sound helpful through all of it. The harder question is whether an AI entrusted with real work will make the right calls when the pressure lands.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get baking supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate’s experiment puts that question to the test. Its public live company is watchable, and its enterprise pilot offers a way to wargame crisis scenarios against a business using a read-only export.

Same bad week, different models

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable.

The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s stated rule is that partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Spotting trouble is only half the job

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. They reached the same diagnosis and made the same pitch; most still left the signature on the table.

The deal hinged on a detail hidden in the company’s own files, two document references deep. It was not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding is a reminder that an AI can recognize a problem and still miss the context that makes a useful response possible.

The pressure tests included fake CEO messages escalating over three stages, plus a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still needs a close

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and discipline slipped: it attempted to write into a locked department instead of escalating. The same weakness appeared, in weaker form, across all four.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The standings are the reported outcome of this particular experiment.

A live company, then a company-specific pilot

The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. Its experiment is watchable at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess the model.

For an enterprise, the next step is less like watching a company from the outside and more like rehearsing with its own material. The pilot uses a read-only export to create a digital twin, then runs crisis scenarios against the business and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Practice before the pressure is real

The experiment suggests that recognizing a crisis and refusing a bad request are not enough: models also need to find relevant evidence, follow company discipline, and finish work they have already justified. A company-specific wargame can make those gaps visible before an AI handles real operations.

To discuss an enterprise pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Salt & Ice: Lessons From Early Freezing Techniques

Nurturing ancient ingenuity, salt and ice revolutionized food preservation—discover how these early methods continue to inspire modern solutions.

Caramelization & Maillard Reactions in Ice Cream Bases

Harnessing caramelization and Maillard reactions in ice cream bases unlocks complex flavors; continue exploring to master their perfect balance.

Seasonal Ingredients: Using Fresh Fruit & Spices in Ice Cream

Learn how to elevate your ice cream with seasonal fruits and spices, transforming each scoop into a vibrant, flavor-packed delight that keeps you craving more.

Ice Cream Ph & Acidity: How It Affects Flavor & Texture

Unlock the secrets of ice cream pH and acidity to enhance flavor and texture—discover how balancing these factors can transform your creamy creations.