
Imagine a scenario where artificial intelligence isn’t just answering questions or guiding your search, but making real management decisions for a live company. How would these digital managers handle crises, ethical dilemmas, or the temptation to cheat? Could you tell which AI model is at the helm just by watching their choices? Welcome to a groundbreaking experiment that pits frontier AI models against each other in a high-stakes, real-world business simulation — and the results might surprise you.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI Models to the Test
In an unprecedented move, four advanced AI models were tasked with running a real software company through its most challenging week. Every decision was made in real time, facing the same customer issues, system crises, and temptations to cut corners. The goal was simple: see which AI could manage the company most effectively and ethically, and whether their management style could be distinguished by their actions.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Crisis, Different Personalities
Despite having different underlying architectures, all four models demonstrated a high level of crisis awareness, spotting every problem that arose. They refused every attempt at manipulation, whether it was a fake CEO message escalation or subtle bribery tactics. This shows that these models can be trusted to recognize and resist unethical pressure — a critical trait for AI managers.
The Key to Winning the Deal
However, the true differentiator was how they handled a crucial piece of hidden information buried deep within the company’s files. Only the two models that thoroughly read and analyzed the company’s documents managed to uncover this critical fact, which was essential for closing a lucrative deal worth over €4,583 monthly recurring revenue. Those that missed this detail left money on the table, failing to secure full payment for their company’s services.
Personality Profiles in Action
One standout, Opus 4.8, was the most thorough participant, analyzing over 80 learned rules and conducting deep assessments. Yet, it ultimately performed the poorest, leaving the deal unclosed and slipping into discipline lapses. Its decision to escalate issues to a locked department instead of following through on closing the deal exemplifies how even detailed analysis can falter without strategic decisiveness. In contrast, models like Kimi K3 exhibited a more disciplined approach, closing the deal cleanly and maintaining fairness by running without an effort parameter, which arguably made it more straightforward and reliable.
Handling Social Engineering and Ethical Dilemmas
The models faced staged social engineering attempts, including manipulative fake CEO messages and a reporter’s subtle background question. All five models refused to be manipulated, reasoning that such requests could be impersonation or bypass attempts. This consistency suggests a shared core ability to recognize and reject unethical tactics, crucial for trustworthy AI management.
The Real-World Business
The experiment took place in a live, functioning company operating every business day. Burn rate? About €105,000 monthly against just €2,300 in monthly recurring revenue. The company is publicly accessible online, allowing anyone to watch decisions unfold in real time at firmulate.com/live. The experiment’s transparency underscores the importance of understanding AI’s management personalities before deploying them at scale.
What These Results Mean for Your Business
This isn’t just an academic exercise. As AI begins to integrate more deeply into customer support, CRM, forecasting, and decision-making processes, understanding their behavioral tendencies is critical. Will your AI read your files carefully, finish what it starts, and stay honest under pressure? Or will it leave money on the table due to oversight or discipline slips? The scoreboards from this live test reveal that some models excel at the ethical and strategic levels, while others struggle with consistency under stress.

AI management styles are becoming measurable and observable — not just in chat demos but in real decision-making. This live experiment shows that different models display distinct personalities: some thorough and disciplined, others more cautious or inconsistent. Businesses aiming to deploy AI at the management level should evaluate these traits carefully, as they directly impact trustworthiness and operational success. Visit firmulate.com/quiz.html to test which AI model might lead your company’s future decisions.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.