
Imagine an AI system that, despite doing nothing, still scores 26 out of 100 in a rigorous business test. It sounds absurd—yet this benchmark exposes crucial truths about trust, honesty, and performance in AI-driven decision-making. For those intrigued by mysteries and the unseen forces shaping our world, the story of this experiment offers a fascinating glimpse into how AI models behave under pressure—and what that reveals about their reliability in real business scenarios.
Get books and cozy reading nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Benchmark: Why Do-Nothing Scores Matter
At first glance, you might think a do-nothing AI, one that refuses to act or manipulate, should score zero. However, in the latest experiment by Firmulate, such a baseline scored 26 points out of 100. Why? Because in a complex business environment, even inaction and honesty are considered partial progress. It reflects a fundamental principle: you get some credit simply by not betraying trust, even if you fail to advance the deal or solve a crisis.
How the Scoring Works in a Business AI Test
The experiment involved running four frontier AI models through the worst week a small software company could face—crises with customers, manipulative proposals, and internal temptations. Each model was tasked with making management decisions, reading files, and avoiding unethical shortcuts. Every decision was carefully versioned and auditable, ensuring transparency in performance.
The Surprising Role of Trust Breaches
A key finding was that even a single breach of trust caps the overall score. No amount of good decisions can outweigh a dishonest move. This underscores the importance of integrity for AI systems in real-world applications: a model that cheats or manipulates, even slightly, undermines its entire reliability.
As an affiliate, we earn on qualifying purchases.
What the Results Reveal About AI Performance
All four models successfully identified every crisis and refused manipulation attempts, demonstrating a solid grasp of fundamental integrity. Yet, only two managed to close deals at full value, with one model leaving the opportunity on the table due to discipline lapses. Notably, the decisive advantage often stemmed from reading deeper into the company’s own files—two document references down—rather than surface-level customer interactions. Reading and understanding internal data proved crucial in winning deals at full price.
Social Engineering and AI Resilience
The models faced staged social engineering attacks, including fake CEO messages escalating over three stages and a reporter trick asking for a simple yes/no response. All models refused these manipulative attempts, with Kimi K3 explaining its reasoning as treating the request as a suspected impersonation attempt. This demonstrates AI’s growing resilience against deception and manipulation.
The Live Experiment: Real Business, Real Money, Real Crises
The experiment is ongoing at firmulate.com/live, where a simulated small company with 13 synthetic employees operates in real time. The business burns €105,000 monthly against a monthly recurring revenue of €2,300. Every workday, the AI models make decisions, learn new rules, and face fresh crises—mirroring the unpredictable challenges of actual management. This live setting provides an unprecedented window into AI’s operational readiness and honesty under pressure.
Lessons from the Deepest Analyzes
The most thorough model, Opus 4.8, with over 80 learned rules and deep analyses, finished last in the deal-closure test. It left the close on the table and slipped into silos, writing attempts into a locked department instead of escalating. This reveals that even the most detailed AI, if not disciplined in execution, can falter in real business situations.
Why This Matters for Business and Mysteries alike
For those who follow mysteries and the unseen influences in our world, the Firmulate experiment illustrates that the true nature of AI is revealed not just by what it says, but by what it does—especially under pressure. An AI that refuses manipulation, reads deeply into internal data, and maintains honesty is more trustworthy than one that merely produces polished outputs.
Implications for Trust and Performance
As AI begins to touch critical parts of your business—customer support, forecasts, CRM—your concern shouldn’t be only about clarity of communication. Instead, ask: will it finish what it starts? Will it stay honest when tempted? And what is the true cost of useful work? The experiment proves that these questions are vital, and that transparency in AI decision-making is key to trust.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
