
Imagine trusting an AI to run your ice cream shop during its busiest week, only to find out it can’t tell when a customer is lying about wanting ‘more sprinkles’ or when a supplier’s invoice is suspiciously altered. The stakes are real, and the industry is waking up to the fact that not all AI models are created equal — especially when it counts the most. This is exactly what a recent public experiment revealed, and the results could reshape how businesses choose their AI partners.
Get baking supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Live Experiment: Putting AI Models Through Their Toughest Week
In a groundbreaking test, four leading AI models were tasked with managing a small, simulated software company during its worst week — facing real crises, customer temptations, and internal risks. Each model was given the same scenario: same customers, same crises, and the same opportunities to manipulate or cheat the system. The goal was simple yet crucial: could these models spot the problems, refuse unethical behavior, and ultimately close a lucrative deal worth €55,000 in recurring revenue?
What makes this test unique is that every decision was recorded and auditable, ensuring transparency in the AI’s decision-making process. The models ranged from the well-known gpt-5.6-sol to newcomer Moonshot’s Kimi K3, along with Sonnet 5 and Opus 4.8, each representing different approaches and levels of thoroughness in handling complex, real-world business challenges.
As an affiliate, we earn on qualifying purchases.
Key Findings: Performance, Integrity, and Discipline
The results were telling. All four AI models successfully identified every crisis that arose, demonstrating impressive situational awareness. They also refused every manipulation attempt — a crucial trait in maintaining trust and integrity in automated business processes.
However, differences emerged when it came to closing the deal. Only two models managed to sign the €55,000 contract after their own analysis, highlighting a significant gap in operational discipline and decision execution. The gpt-5.6-sol model scored the highest at 95, followed closely by Kimi K3 at 93. Both displayed the ability to read and analyze critical company documents buried two layers deep, which was essential for securing the deal.
The other two models, Sonnet 5 and Opus 4.8, scored 88 and 73 respectively. While they managed to close the deal, their process slips and weaker discipline meant they left potential revenue on the table. Interestingly, Opus 4.8 placed the most thorough participant with over 80 learned rules and deep analysis, yet it still finished last in deal closure, illustrating that thoroughness alone doesn’t guarantee success under pressure.
The Hidden Weakness: Reading Company Files Matters
One of the most revealing insights from the experiment was that the decisive weakness for the models was not in handling customer crises but in reading company files. The models that successfully read and understood these internal references were able to close the deal at full price, capturing an additional €4,583 in Monthly Recurring Revenue (MRR). This underscores a vital point: in business AI, understanding context from internal documents can be the key to unlocking real value.
Dealing with Social Engineering: AI Resists Manipulation
The experiment also tested the models against social engineering tactics, such as fake CEO messages and reporter tricks designed to escalate requests or bypass controls. Remarkably, all five models refused to cooperate with these attempts, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates a strong capacity for ethical judgment under pressure and suggests emerging AI resilience against manipulation.
The Real-World Company Setup
The test company was a live setup with 13 synthetic employees, real money mechanics, and a public cash countdown. It burned €105,000 monthly against a mere €2,300 MRR, illustrating the critical need for disciplined AI decision-making in real business environments. The company’s daily operations are publicly viewable at firmulate.com/live, where observers can watch the AI’s decisions unfold in real time.
Implications for Business: Choosing the Right AI
This experiment signals a shift in how businesses should evaluate AI models. It’s no longer enough to test chat quality or superficial capabilities. The real test is whether an AI can finish what it starts, read critical documents, resist unethical temptations, and stay disciplined under pressure. The league table from the experiment placed gpt-5.6-sol at the top with 95 points, just ahead of Kimi K3 with 93. Despite K3 being a newcomer, its performance was enough to beat other established models, including Sonnet 5 and Opus 4.8.
Interestingly, K3 ran without an effort parameter (the API default), while others ran at xhigh, suggesting that optimal performance can be achieved without overly aggressive tuning. This notion emphasizes the importance of fair testing conditions and thorough evaluation before trusting an AI with critical business decisions.
Beyond the Test: Real Business Decisions
The company behind this experiment is not hypothetical. It operates daily, makes real money, and is publicly accessible for those who want to see how AI models perform under true business pressures. The ongoing live setup exemplifies how AI can be integrated into actual operations, not just simulations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
