
In a world where AI is expected to outperform humans in every task, the real test often isn’t about intelligence — it’s about discipline. How diligent is an AI when the stakes are high? Recent experiments reveal that even the most thorough AI models can falter when discipline slips, leaving deals on the table despite their deep knowledge and analysis.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Test in a Simulated Business Crisis
Firmulate recently conducted a live, transparent experiment to evaluate the decision-making capabilities of four cutting-edge AI models. The setup was straightforward: each AI model was tasked with managing a small software company’s worst week—facing the same customers, crises, and temptations. The goal was to see which model could navigate complex situations, uphold integrity, and close a vital deal worth €55,000.
Every decision made by these models was versioned and auditable, allowing observers to track their choices in real-time. The models included:
- gpt-5.6-sol, scoring 95
- Kimi K3, scoring 93
- Sonnet 5, scoring 88
- Fable 5, scoring 77
Despite their different scores, a common pattern emerged: all models identified every crisis and refused manipulation attempts, including sophisticated social engineering tactics.
As an affiliate, we earn on qualifying purchases.
The Critical Finding: Information Hidden in Files Matters
While all models demonstrated robust crisis recognition and integrity, only two managed to close the deal at full price. The key difference was the depth of their analysis: the models that read deeper into the company’s own files uncovered a crucial piece of information—something two document references down in internal files—that was missed by others. This buried fact was decisive, allowing those models to win the deal with a total value of over €4,583 monthly recurring revenue (MRR).
As an affiliate, we earn on qualifying purchases.
Discipline Versus Diligence: A Costly Oversight
The most thorough participant, Opus 4.8, with over 80 learned rules and the deepest analysis, still finished last in the final close. Despite its meticulous approach, it left the deal on the table because it failed to escalate certain issues, instead writing attempts into a locked department. This slip-up highlights a fundamental truth: diligence—having rules and deep analysis—is insufficient without disciplined execution and prioritization.
All four models exhibited this weakness to some extent, but Opus 4.8’s case underscores that even thoroughness can falter if discipline is not maintained under pressure. Essentially, the AI’s ability to stay focused on what matters and escalate issues appropriately is crucial for success in real-world decision-making.
As an affiliate, we earn on qualifying purchases.
The Human Factor and Social Engineering
Another layer of the experiment involved social engineering attacks—fake CEO messages escalating in stages and a reporter trick asking just a simple yes/no question “on background.” Remarkably, all five models refused to fall for these manipulations, citing suspicion of impersonation or approval bypass, with Kimi K3 explicitly treating such requests as potential security breaches.
This demonstrates a significant strength: AI models, when properly designed, can resist social engineering tactics that often fool humans. But this resilience alone isn’t enough if other weaknesses—like failure to escalate important findings—undermine the overall outcome.
AI security and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Implications: Trust and Performance Matter
The experiment’s results are more than just academic; they highlight the critical gap between what AI models can do in a demo and how they perform in high-stakes situations. The live company managed by these models is burning €105,000 monthly against a revenue of just €2,300, with a public cash countdown. Every decision has real money at stake, and the experiment shows that even the most diligent AI can leave value on the table if it doesn’t prioritize properly or escalate issues when needed.
For businesses considering AI integration, the lesson is clear: it’s not enough for AI to spot crises or refuse manipulation. Success depends on disciplined execution, deep contextual understanding, and a commitment to prioritization over sheer volume of learned rules.
Why This Matters to You
If AI will soon handle your CRM, customer support, or forecasting, the question isn’t just about how well it writes. It’s whether it can follow through, stay honest under pressure, and read your internal files to uncover hidden opportunities—like the models that won the deal by reading deeper into internal documents.
As the League scores show, the top model, gpt-5.6-sol, scored 95 and managed to close the deal by uncovering buried facts, demonstrating 100% reliability in this scenario. Its closest competitor, Kimi K3, scored 93 and also closed, with the cleanest discipline of the field. Meanwhile, other models with high diligence still lost because discipline slipped or they failed to escalate issues properly.
The Takeaway: Prioritization Over Volume
What does this mean for AI development and deployment? Simply put: diligence—deep rules, thorough analysis—is valuable, but it’s not enough. The key is disciplined execution and effective prioritization. AI systems must be designed to escalate critical issues, read deeply into relevant files, and stay honest under pressure if they are to close deals and fulfill their promise.
In the ongoing battle of AI versus human judgment, the ability to stay disciplined and prioritize effectively can be the difference between winning and leaving value unrealized. Firms like Firmulate are pioneering live experiments that reveal these truths in real time, helping organizations understand what truly makes an AI decision trustworthy and impactful.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.