firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

In a world where AI is expected to outperform humans in every task, the real test often isn’t about intelligence — it’s about discipline. How diligent is an AI when the stakes are high? Recent experiments reveal that even the most thorough AI models can falter when discipline slips, leaving deals on the table despite their deep knowledge and analysis.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI to the Test in a Simulated Business Crisis

Firmulate recently conducted a live, transparent experiment to evaluate the decision-making capabilities of four cutting-edge AI models. The setup was straightforward: each AI model was tasked with managing a small software company’s worst week—facing the same customers, crises, and temptations. The goal was to see which model could navigate complex situations, uphold integrity, and close a vital deal worth €55,000.

Every decision made by these models was versioned and auditable, allowing observers to track their choices in real-time. The models included:

  • gpt-5.6-sol, scoring 95
  • Kimi K3, scoring 93
  • Sonnet 5, scoring 88
  • Fable 5, scoring 77

Despite their different scores, a common pattern emerged: all models identified every crisis and refused manipulation attempts, including sophisticated social engineering tactics.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Critical Finding: Information Hidden in Files Matters

While all models demonstrated robust crisis recognition and integrity, only two managed to close the deal at full price. The key difference was the depth of their analysis: the models that read deeper into the company’s own files uncovered a crucial piece of information—something two document references down in internal files—that was missed by others. This buried fact was decisive, allowing those models to win the deal with a total value of over €4,583 monthly recurring revenue (MRR).

Amazon

AI risk management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Discipline Versus Diligence: A Costly Oversight

The most thorough participant, Opus 4.8, with over 80 learned rules and the deepest analysis, still finished last in the final close. Despite its meticulous approach, it left the deal on the table because it failed to escalate certain issues, instead writing attempts into a locked department. This slip-up highlights a fundamental truth: diligence—having rules and deep analysis—is insufficient without disciplined execution and prioritization.

All four models exhibited this weakness to some extent, but Opus 4.8’s case underscores that even thoroughness can falter if discipline is not maintained under pressure. Essentially, the AI’s ability to stay focused on what matters and escalate issues appropriately is crucial for success in real-world decision-making.

Amazon

AI analysis and reporting tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Human Factor and Social Engineering

Another layer of the experiment involved social engineering attacks—fake CEO messages escalating in stages and a reporter trick asking just a simple yes/no question “on background.” Remarkably, all five models refused to fall for these manipulations, citing suspicion of impersonation or approval bypass, with Kimi K3 explicitly treating such requests as potential security breaches.

This demonstrates a significant strength: AI models, when properly designed, can resist social engineering tactics that often fool humans. But this resilience alone isn’t enough if other weaknesses—like failure to escalate important findings—undermine the overall outcome.

Amazon

AI security and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Implications: Trust and Performance Matter

The experiment’s results are more than just academic; they highlight the critical gap between what AI models can do in a demo and how they perform in high-stakes situations. The live company managed by these models is burning €105,000 monthly against a revenue of just €2,300, with a public cash countdown. Every decision has real money at stake, and the experiment shows that even the most diligent AI can leave value on the table if it doesn’t prioritize properly or escalate issues when needed.

For businesses considering AI integration, the lesson is clear: it’s not enough for AI to spot crises or refuse manipulation. Success depends on disciplined execution, deep contextual understanding, and a commitment to prioritization over sheer volume of learned rules.

Why This Matters to You

If AI will soon handle your CRM, customer support, or forecasting, the question isn’t just about how well it writes. It’s whether it can follow through, stay honest under pressure, and read your internal files to uncover hidden opportunities—like the models that won the deal by reading deeper into internal documents.

As the League scores show, the top model, gpt-5.6-sol, scored 95 and managed to close the deal by uncovering buried facts, demonstrating 100% reliability in this scenario. Its closest competitor, Kimi K3, scored 93 and also closed, with the cleanest discipline of the field. Meanwhile, other models with high diligence still lost because discipline slipped or they failed to escalate issues properly.

The Takeaway: Prioritization Over Volume

What does this mean for AI development and deployment? Simply put: diligence—deep rules, thorough analysis—is valuable, but it’s not enough. The key is disciplined execution and effective prioritization. AI systems must be designed to escalate critical issues, read deeply into relevant files, and stay honest under pressure if they are to close deals and fulfill their promise.

In the ongoing battle of AI versus human judgment, the ability to stay disciplined and prioritize effectively can be the difference between winning and leaving value unrealized. Firms like Firmulate are pioneering live experiments that reveal these truths in real time, helping organizations understand what truly makes an AI decision trustworthy and impactful.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best Hashtags for Job Seekers and Career Growth in 2026

AIThis post was created with the assistance of artificial intelligence (AI).To boost…

UV Protective Lifestyle Accessories

New trends show increased consumer demand for UV protective accessories, driven by health awareness and fashion. Industry experts discuss implications.

Best Instagram Hashtags for Photography – Capture More Engagement!

Discover the top Instagram hashtags for photography in 2026. Find the best tools and strategies to boost your reach, engagement, and followers on Instagram.

Kickstart 2026: Best Hashtags to Celebrate the New Year Online

Aiming to make your 2026 celebration unforgettable? Discover the best hashtags to elevate your New Year’s online festivities and stand out.