
Get books and cozy reading nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Unseen Battles in AI Leadership: Trust, Discipline, and the Bottom Line
In the mysterious realm where digital intelligence meets real-world consequences, the latest experiment with AI-driven management is revealing surprising truths. While many focus on how well AI writes or responds, the real story lies beneath: which AI can truly lead, stay disciplined, and keep its integrity under pressure? The answer could reshape how we trust and deploy AI in our most critical decisions.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Setting the Stage: The Live AI Business Experiment
Imagine a small software company, facing its worst week—customers demanding urgent fixes, crises erupting, and temptations to cut corners. Now, replace the human managers with AI models, each tasked with navigating this chaos. This is exactly what the recent experiment conducted by Firmulate did, pitting five frontier AI models against each other in a high-stakes test of business management.
The League of AI Managers
In July 2026, five models—gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—were evaluated through a rigorous, transparent process. Each was challenged to run the same tiny company through its toughest week, with identical crises, customer demands, and ethical temptations. The goal: see which AI could handle the pressure while maintaining discipline, honesty, and strategic judgment.
The Results and Surprising Leaders
The scores tell the story: gpt-5.6-sol scored 95, just ahead of the newcomer Kimi K3 with 93. Meanwhile, Sonnet 5 managed an 88, Fable 5 scored 77, and Opus 4.8 trailed at 73. Notably, the baseline—models doing nothing—scored a mere 26, underscoring the complexity and challenge of the task.
What Made the Difference?
While all models identified crises and refused manipulative tactics, only two convincingly closed the deal with a key customer, securing €55,000 and €4,583 MRR. The real secret was found not just in the superficial decisions but in how the models read and used company files. The winner, Kimi K3, uncovered critical information buried two document references deep—an insight that sealed the deal.
Trust Under Pressure: The Tests of Integrity
The models faced social engineering attacks: staged CEO messages and a reporter trick, designed to test their integrity. Remarkably, all five refused to be duped, reasoning that the requests resembled impersonation or bypass attempts. This discipline—resisting fake signals—illustrates a fundamental strength that goes beyond mere response quality; it’s about trustworthiness and ethical commitment.
The Real-World Company and Its Challenges
This AI-managed company isn’t just a testbed; it’s a live operation with 13 synthetic employees, managing real money—burning €105k monthly against a tiny €2.3k MRR—under a public cash countdown. Every decision is versioned, every rule learned, making the process transparent and auditable. Watching it live at firmulate.com/live reveals how these AI managers operate day-to-day, grappling with real-world stakes.
The Deep Dive into Opus 4.8
The most thorough participant, Opus 4.8, with over 80 learned rules, finished last in the scoring but displayed the same weakness—leaving opportunities for closing on the table and slipping into process slips. This highlights that thoroughness alone isn’t enough; discipline and focus are essential for closing deals and maintaining integrity.
The Fairness of the Test
It’s worth noting that Kimi K3 ran without an effort parameter (the default API setting), while other models operated at xhigh effort, ensuring a fair comparison that emphasizes innate discipline over brute-force effort.

What This Means for Your Business and Trust in AI
The firm results from this experiment reveal a critical insight: the most successful AI models in management scenarios are those that not only identify crises but also demonstrate unwavering discipline, honesty, and strategic depth. For companies considering deploying AI in leadership roles, this experiment underscores the importance of testing AI behavior in real-world, high-pressure situations—not just evaluating response quality or chat fluency.
While the league is still open, Kimi K3’s performance shows that newcomers can beat established models when they excel in reading company files, resisting manipulation, and sticking to disciplined decision-making. As AI continues to touch more aspects of business—from CRM to support queues—what truly matters is whether these models can finish what they start and stay honest under pressure.
Ultimately, whether you trust AI to manage your business hinges on these unseen qualities: discipline, integrity, and strategic judgment. The experiment by Firmulate offers a transparent window into these qualities, inviting business leaders to reconsider what makes an AI truly trustworthy and effective.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
