
Imagine a company with no employees, burning through €105,000 each month, yet still publicly battling to survive. It’s not a science fiction scenario, but a real-time experiment demonstrating the future of AI-driven management. This is the story of Firmulate, a live company run entirely by AI models, constantly tested in a high-stakes environment that reveals both their strengths and vulnerabilities.
The Living Experiment: A Company Without People
At the heart of this experiment is a small software business simulated with 13 synthetic employees, each driven by advanced AI models. Unlike most AI demos, this setup is fully operational, with real money mechanics at play. The company faces ongoing crises, customer negotiations, and internal decision-making — all while being publicly accessible at firmulate.com/live.html.
What makes this project extraordinary is its transparency and rigor. Every workday, the AI models make decisions, and these are versioned and auditable. The entire operation is designed to test whether these models can handle the complexities of real management: detecting crises, resisting manipulation, and making profitable deals without human oversight.

Construction Program Management – Decision Making and Optimization Techniques
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Challenges: Crises, Manipulations, and Ethical Tests
During a recent experiment, four frontier AI models — including GPT-5.6-sol and Kimi K3 — were subjected to the same simulated week of chaos. They faced the same crises, customer demands, and ethical dilemmas, such as social engineering attempts to manipulate the AI into bypassing protocols. Remarkably, all four models identified every crisis and refused every attempt at manipulation. They demonstrated integrity under pressure, refusing to sign a €55,000 deal that their own analysis recommended.
Digging deeper, the models’ ability to uncover critical information was tested. The decisive weakness wasn’t in the customer interactions but buried in the company’s internal files. Only those models that read and understood the internal documents were able to close the deal at full price, adding €4,583 to monthly recurring revenue (MRR). This emphasizes that reading and understanding internal data can make or break business outcomes in AI management.
Real Money, Real Consequences
The company’s cash flow paints a stark picture: it burns €105,000 monthly against a tiny €2,300 MRR, with a public countdown looming. Every decision, every rule learned, and every version of the decision-making process is visible and tracked, creating a transparent, dynamic laboratory for AI and management experts alike.

The Scalpel and the Algorithm: Reclaiming Ethical Clarity and Clinical Confidence in the Age of AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Results Tell Us
Among the models tested, Opus 4.8, which had the most thorough analysis with over 80 learned rules, ended up in last place — initially leaving the deal on the table and slipping into operational slips like writing attempts into a locked department instead of escalating. Conversely, Kimi K3, running without an effort parameter, performed better and closed the deal. The results underline that discipline, reading internal data, and adherence to rules are critical factors, often more than raw power or complexity.
These findings are part of a broader assessment called the Crucible League, where models are scored from 26 (bottom) to 95 (top) based on their performance in these simulated crises. The top scorer, GPT-5.6-sol, achieved a 95, having uncovered critical internal facts and closed the deal at full value.

How AI Agents Work: Tools, Memory, and Autonomous Decision-Making (The AI Security & Hacking Bible: Protect and Exploit LLMs and Autonomous Agents)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for the Future of Business
This experiment is more than a tech showcase; it’s a glimpse into the future of AI in real-world management. As companies increasingly integrate AI agents into their operations — from customer support to strategic decision-making — the key questions aren’t just about how well they write but whether they are trustworthy, disciplined, and capable of finishing what they start.
The current setup allows enterprise leaders to run “wargames” against their own business models without risking real systems, at firmulate.com/pilot.html. This offers a safe way to test how AI might behave under stress and whether it can be relied upon to act ethically and effectively under pressure.

Information Systems for Crisis Response and Management in Mediterranean Countries: 4th International Conference, ISCRAM-med 2017, Xanthi, Greece, … in Business Information Processing, 301)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Final Thoughts: Watching AI’s Moral Muscles in Action
Firmulate’s live experiment is a provocative, transparent window into the evolving relationship between humans and AI in management. It exposes not only the technical capabilities but also the ethical stamina of these models. As the company’s cash dwindles and its future hangs in the balance, it provides a real-time case study of AI’s potential and limitations in high-stakes decision environments.

Firmulate’s public, real-time management test reveals AI’s ability to handle crises, resist manipulation, and make profitable decisions—showing the future of trustworthy AI in business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html