
Imagine a company with no human employees, yet constantly battling to stay afloat—while every decision it makes is publicly observable, auditable, and driven by AI models. Welcome to the frontier of build-in-public AI experimentation, where every workday is a live demonstration of the strengths and weaknesses of AI in management.
The Real-Time Experiment in AI-Driven Management
At the heart of this bold experiment is a small software company run entirely by artificial intelligence models, with no human workers involved. Every decision, crisis response, and negotiation is made by AI, tested against real-world scenarios that mimic the worst week of a typical business. The company burns through €105,000 each month but earns only €2,300 in recurring revenue, making its survival a daily challenge watched by all.
The Live Company and Its Mechanics
This company isn’t fictional. It’s a live, evolving setup hosted at firmulate.com/live.html, where anyone can see its daily operations. With 13 synthetic employees—each represented by a sophisticated AI model—the operation is a window into the future of automation and management intelligence. The company operates with over 680 self-learned rules and every decision is versioned and recorded, allowing observers to trace how each AI model responds to crises and temptations.
Testing the Limits of AI Judgment
Four frontier AI models were challenged to navigate the same set of crises, customer interactions, and ethical dilemmas. The scenario was rigorous: all models faced identical crises, and their performance was measured based on their ability to identify opportunities, refuse manipulation, and ultimately close deals. The results were enlightening — all four AI models successfully identified every crisis and refused every manipulation attempt. But only two managed to close a deal at full price, which was worth over €4,500 in monthly revenue.
The Hidden Weaknesses and the Power of Data
The decisive factor wasn’t in the obvious decisions. The models that succeeded had an edge in reading deeper into internal documents—data buried two document references within the company’s own files. Those who read the full context closed the deal at full value, proving that understanding internal context is crucial for AI to make profitable decisions in complex environments.
Resisting Social Engineering and Ethical Challenges
In a series of staged social engineering tests, AI models faced fake CEO messages escalating in urgency and a reporter’s subtle tricks, where they were asked for quick approvals or background information. Remarkably, all five models refused to be manipulated, citing reasons like suspicion of impersonation or bypassing protocols. This underscores the importance of ethical safeguards embedded within these AI systems.
The Cost of Running a Zero-Employee Business
The entire operation runs with no human staff, yet it’s bleeding money — burning €105,000 per month against a small income stream of €2,300. It’s a stark reminder that automation isn’t about cost savings alone; it’s about understanding how AI can manage crises, negotiations, and integrity in real time, even if the business isn’t profitable yet.
Insights from the Top-Performing Model
The most thorough AI—called Opus 4.8—had over 80 learned rules and conducted deep analysis. Despite its discipline, it left a potential close on the table and faltered in escalating issues properly. Interestingly, even the best models demonstrate weaknesses, highlighting that AI management at this level requires continuous refinement and oversight.
What This Means for Business Decision-Making
This experiment isn’t just a tech showcase; it’s a mirror for real companies contemplating AI integration. What matters isn’t just whether AI can generate good-looking chat responses, but whether it can finish what it starts, read critical internal data, and remain honest under pressure. These are the qualities that determine whether AI will help or hinder your enterprise.

The AI Operating System: A Field Guide to Working With Intelligence, Not Under It
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Build-in-Public — Transparency as the Future
The entire setup is openly accessible and versioned every workday, inviting anyone to watch the decision-making process unfold in real time. This transparency fosters trust and helps developers and business leaders understand AI’s true capabilities and limitations in managing complex, real-world scenarios.
Beyond the Experiment: Practical Tools and Engagement
For businesses interested in testing their own AI workforce, the platform offers read-only simulations of their operations, allowing leaders to run wargames without risking real systems or data. More information can be found at firmulate.com/pilot.html.

This live AI management experiment vividly demonstrates that AI can identify crises, refuse manipulation, and close deals—yet still exhibits weaknesses that reveal where human judgment remains critical. Its transparency and continuous decision recording make it a unique window into the future of AI-led management, offering valuable lessons for any enterprise considering automation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

AI Incident Response Systems: Crisis Management AI | AI Security Playbooks | Digital Forensics Enhanced | AI-Driven Incident Management | AI Forensic Innovations | Automated Security Solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI negotiation simulation platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Safeguards in a World of Ambient Intelligence (The International Library of Ethics, Law and Technology, 1)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.