
Imagine an AI that doesn’t just chatter or generate text—it actually reads and understands your internal files before making a decision. In the world of business, this ability can be the difference between sealing a deal or losing it. Now, recent experiments with artificial intelligence models show that reading your documents deeply is a measurable, decisive factor in securing high-stakes agreements.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Hidden Depths of AI Decision-Making
In a groundbreaking live experiment, four leading AI models were tasked with running a small software company through its most challenging week. The scenario was realistic: same customers, same crises, same temptations to cut corners or manipulate. Every decision made by these models was recorded, versioned, and auditable, creating a transparent window into their reasoning processes.
The results? All four models identified every crisis and refused every manipulation attempt—showing strong ethical behavior and crisis management. However, only two of these models managed to close a critical €55,000 deal, earning a full signature based on their own analysis. The other two, despite similar diagnoses, left the deal on the table, missing a key opportunity.
As an affiliate, we earn on qualifying purchases.
The Crucial, Buried Fact
What made the difference? The decisive weakness was buried two document references deep within the company’s files—information that wasn’t visible on the surface or in the initial customer interactions. The models that delved into the internal documents, reading thoroughly and understanding context, were able to uncover this hidden detail and ultimately win the deal.
This finding underscores a vital property: for AI to be truly effective in business, it must read and interpret the internal files before answering or making decisions. Superficial or surface-level analysis isn’t enough in high-stakes negotiations.
Beyond Chat: Measuring True Business Utility
The experiment highlights that the real value of AI in business isn’t just in generating convincing conversations or responses. Instead, it’s in its ability to finish what it starts, to read your files thoroughly, and stay honest under pressure. For example, during a social engineering test, where fake CEO messages escalated in complexity, all models refused manipulation attempts. Kimi K3 explained its refusal as a suspicion of impersonation, demonstrating an understanding of the importance of verification and trust.
In practice, this means AI systems should be evaluated not only on their chat quality but on their capacity to read, comprehend, and act on internal data. The models that excel in these areas can uncover critical hidden truths, make more accurate decisions, and close deals more reliably.
The Live Experiment: Testing in a Real-World-Like Setting
The live experiment runs at firmulate.com/live and simulates a small company’s daily operations, including real money mechanics and self-learned rules. It leverages a synthetic company with 13 employees, burning €105,000 per month against a €2,300 monthly recurring revenue, illustrating the importance of disciplined decision-making. The experiment is transparent, with every move versioned and available for review.
Within this environment, four frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—are tested on their ability to navigate crises, resist manipulation, and close deals. Notably, the most thorough model, Opus 4.8, with over 80 learned rules and deep analysis, finished in last place—illustrating that deeper analysis alone doesn’t guarantee success if process discipline slips or opportunities are missed.
The Takeaway for Business Leaders
What does this mean for companies integrating AI? The key property isn’t just language fluency or chat prowess. It’s the AI’s capacity for comprehensive reading and understanding of internal data before acting. This ability can be the difference between winning or losing a significant deal, especially when crucial information is buried in documents.
As one of the models, Kimi K3, noted: ‘Treat the request as a suspected approval-bypass / possible impersonation.’ This cautious, trust-aware approach exemplifies how AI can act responsibly in sensitive situations, especially when reading internal files deeply.
Empowering Your Business with Better AI Testing
Enterprises eager to harness this capability can simulate their own operations against a read-only export of their data, testing how their AI might perform in real negotiations or crisis management. The platform at firmulate.com/pilot.html offers such wargaming environments, ensuring your AI is ready before deployment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
College move-in / dorm season Picks
dorm essentials
As an affiliate, we earn on qualifying purchases.