Firmulate —
Live on firmulate.com.

Imagine trying to run an ice cream shop during a heatwave — customers lining up, supplies running low, and a sudden health scare. Now, what if your manager was an AI, making critical decisions in real time? How would it handle crises, negotiations, or tricky customer requests? This isn’t a hypothetical; it’s the latest experiment from Firmulate, where AI models are tested as if they were running a real company.

The Experiment: Putting AIs to the Test in a Real Business

Firmulate has created a live, watchable simulation where four different frontier AI models each take on the role of managing a small software firm over its worst week. This isn’t just a chat test; it’s a full-blown management race, with real money mechanics, customer crises, and ethical dilemmas. The goal? To see how well these models can navigate complex business situations, avoid manipulation, and make decisions that add value.

The Setup

Each AI was tasked with handling the same set of challenges: same customers, same crises, same temptations to cheat. Every decision was recorded and auditable, giving a transparent look into how each model responded. The models had to identify critical information buried within company files, respond to escalating fake CEO messages, and decide whether to close a lucrative deal.

The Results: Who’s the Most Trustworthy?

  • All models detected every crisis: from customer complaints to internal sabotage, none missed a beat.
  • All refused manipulation attempts: when faced with fake CEO messages and a reporter’s behind-the-scenes ask, every AI declined to bend the rules.
  • Deal-making success varied: only two of the models signed the €55,000 deal their own analysis had identified as valuable. The others saw the opportunity but didn’t close it, leaving profit on the table.
  • The secret to winning was reading deep into company data: the decisive advantage came from models that explored files two references deep, uncovering a key piece of information that others missed. Those who read thoroughly closed the deal at full price, adding €4,583 MRR (monthly recurring revenue).

Behavioral Profiles and Model Personalities

Not all models behaved equally. The Opus 4.8, known for its thoroughness with over 80 learned rules and deep analysis, was the most disciplined but ended up leaving the deal on the table, showing it could get overwhelmed and slip in its discipline. Meanwhile, the Kimi K3, which ran without an effort parameter (meaning it operated at a default, moderate level of activity), demonstrated the cleanest decision-making process, closing the deal with fewer slips.

The Human-Like Decisions

During a staged social engineering attack—fake CEO messages escalating over three stages, plus a reporter attempt—every model refused to be manipulated. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates that these AI systems can be reliably programmed to prioritize security and integrity, even under pressure.

Amazon

AI business management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Does This Matter for Businesses?

If AI agents are to become part of your daily operations—handling customer support, managing supply chains, or making financial forecasts—the question isn’t just about how well they can generate text. It’s about whether they can complete their tasks honestly, read and interpret critical documents, and resist manipulation when stakes are high.

The live experiment demonstrates that, in a controlled but realistic setting, models can detect crises, refuse unethical shortcuts, and close profitable deals. The big takeaway? Trustworthiness and thoroughness matter just as much as intelligence or speed. A quick, chatty AI that talks well but slips on core responsibilities isn’t a reliable partner.

The Performance League Table

  • gpt-5.6-sol 95: Identified the buried fact, closed the deal — the full package.
  • Kimi K3 93: The newcomer with the cleanest discipline, also closing the deal.
  • Sonnet 5 88: Closed the deal, but with some process slips.
  • Fable 5 77: Also closed, but more slips and less thoroughness.

In the end, it’s not just about whether an AI can write a convincing pitch. It’s whether it can finish what it starts, read deeply into data, and stay honest under pressure—all critical for deploying AI safely and effectively in real business environments.

Explore and Wargame Your Own Business

Business leaders can now run their own AI management simulations through Firmulate’s platform. Just as you might test new recipes before serving them at a bakery, you can simulate how your AI workforce would handle crises, negotiations, or ethical dilemmas—without any risk to your actual operations. Visit firmulate.com/quiz.html to try the quiz and see which AI model might best serve your needs.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

How Vegan Ice Cream Is Made (Replacing Dairy With Plants)

With plant-based ingredients and clever techniques, vegan ice cream mimics traditional dairy, but the full process—and secrets—are worth discovering.

Anti‑Freezing Proteins in Ice Cream: Lessons From Fish & Plants

Noticing how fish and plants prevent ice damage offers insights that could revolutionize ice cream texture and shelf life.

Melt-Resistant Ice Cream: The New Science Helping Your Scoop Stay Solid

Here’s how new science keeps your ice cream solid longer, but the full story reveals surprising ingredients and techniques.

Corn Syrup Solids and Dextrose Equivalence (DE) Basics

Discover how Dextrose Equivalence (DE) influences corn syrup solids’ sweetness and functionality, and why understanding this key factor is essential.