
Imagine running a small, struggling software company where every decision could make or break your future — and doing so with the help of artificial intelligence. How well can these models navigate crises, read between the lines, and keep their integrity under pressure? This is not a sci-fi scenario but a real-world experiment unfolding live, where frontier AI models are put through their paces in a simulated business environment.
Get art and craft supplies delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
The Live Test of AI Management Skills
At the heart of the experiment is a small software firm facing its worst week, complete with difficult customers, crises, and temptations to cut corners. Four advanced AI models — including GPT-5.6-SOL, Kimi K3, Sonnet 5, and Fable 5 — each run the same company through this chaos. Every decision they make is recorded, versioned, and auditable, providing an unprecedented window into how AI handles complex management tasks in real time.
As an affiliate, we earn on qualifying purchases.
The Surprising Performance Results
Despite the high-stakes scenario, all four models demonstrated a fundamental ability: they identified every crisis and refused every attempt at manipulation. This includes social engineering attacks such as staged CEO messages and background verification requests, which all models successfully rejected, citing concerns about impersonation or bypassing approval protocols.
The key difference emerged in their ability to close deals. Only two models, GPT-5.6-SOL and Kimi K3, signed the €55,000 contract their own analysis had earned — a mark of trust and thoroughness. The other two, Sonnet 5 and Fable 5, left the deal on the table, showing gaps in process discipline or strategic judgment. Interestingly, the decisive factor wasn’t just the decision itself but where the models looked for crucial information.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness and the Power of Context
Deep within the company’s files, two document references held the key to securing the deal at full value. Models that read and incorporate this knowledge secured the best outcome, highlighting how access to relevant data influences decision quality. This underscores an important point: AI models that dig beneath surface-level information can achieve significantly better results than those relying solely on immediate inputs.
As an affiliate, we earn on qualifying purchases.
Personality and Decision Styles in AI
Beyond raw performance, the experiment revealed different management personalities. For example, Opus 4.8, the most thorough participant with over 80 learned rules and deep analysis, ultimately slipped in closing discipline, leaving opportunities unseized and failing to escalate issues properly. Conversely, Kimi K3 ran without an effort parameter, emphasizing fairness and straightforwardness, and maintained the discipline necessary to close deals effectively.
As an affiliate, we earn on qualifying purchases.
The Real-World Relevance
All decisions, both good and bad, are happening in a publicly observable business environment. The company, with 13 synthetic employees and real financial mechanics, burns €105,000 monthly against a tiny €2,300 MRR, illustrating how AI management doesn’t just matter in theory but has real fiscal consequences. The experiment runs every business day, with every decision versioned, allowing viewers to see decision-making in action at firmulate.com/live.
What Does It Mean for Your Business?
If AI agents are to become part of your operations—handling customer relationships, support queues, or forecasting—the critical questions aren’t about their language quality but about their integrity and follow-through. Will they finish what they start? Will they read relevant files before acting? Will they stay honest under pressure? These are the factors that determine whether AI will be an asset or a liability.
The League Table and What It Tells Us
- GPT-5.6-SOL scored 95 and found the hidden fact, sealing the deal at full price.
- Kimi K3 scored 93, also closing the deal with the cleanest discipline.
- Sonnet 5 scored 88, with minor slips in process.
- Fable 5 scored 77, leaving the deal unclosed despite recognizing the opportunity.
- The baseline score was 26, underscoring the importance of capability beyond partial progress.
Experience the Live Experiment
Unlike static demonstrations, this is a real company in action, every decision made in a live environment, with results available for scrutiny. You can see it all unfold, read employees’ actual comments, or even run a similar test on your own business data through the dedicated platform at firmulate.com/pilot.html.
Why This Matters for Arts, Crafts & Culture Enthusiasts
Just as artists and artisans understand the importance of craftsmanship, the same principle applies to AI management. Trust, discipline, and attention to detail matter just as much in digital decision-making as in traditional creative work. This experiment shows that AI models can be trained to behave reliably — but only if their decision-making processes are rigorously tested and monitored.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
