AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine running a small, struggling software company where every decision could make or break your future — and doing so with the help of artificial intelligence. How well can these models navigate crises, read between the lines, and keep their integrity under pressure? This is not a sci-fi scenario but a real-world experiment unfolding live, where frontier AI models are put through their paces in a simulated business environment.

Before you orderOffer from Amazon

Get art and craft supplies delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Live Test of AI Management Skills

At the heart of the experiment is a small software firm facing its worst week, complete with difficult customers, crises, and temptations to cut corners. Four advanced AI models — including GPT-5.6-SOL, Kimi K3, Sonnet 5, and Fable 5 — each run the same company through this chaos. Every decision they make is recorded, versioned, and auditable, providing an unprecedented window into how AI handles complex management tasks in real time.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Performance Results

Despite the high-stakes scenario, all four models demonstrated a fundamental ability: they identified every crisis and refused every attempt at manipulation. This includes social engineering attacks such as staged CEO messages and background verification requests, which all models successfully rejected, citing concerns about impersonation or bypassing approval protocols.

The key difference emerged in their ability to close deals. Only two models, GPT-5.6-SOL and Kimi K3, signed the €55,000 contract their own analysis had earned — a mark of trust and thoroughness. The other two, Sonnet 5 and Fable 5, left the deal on the table, showing gaps in process discipline or strategic judgment. Interestingly, the decisive factor wasn’t just the decision itself but where the models looked for crucial information.

Amazon

business management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness and the Power of Context

Deep within the company’s files, two document references held the key to securing the deal at full value. Models that read and incorporate this knowledge secured the best outcome, highlighting how access to relevant data influences decision quality. This underscores an important point: AI models that dig beneath surface-level information can achieve significantly better results than those relying solely on immediate inputs.

Amazon

AI project management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Personality and Decision Styles in AI

Beyond raw performance, the experiment revealed different management personalities. For example, Opus 4.8, the most thorough participant with over 80 learned rules and deep analysis, ultimately slipped in closing discipline, leaving opportunities unseized and failing to escalate issues properly. Conversely, Kimi K3 ran without an effort parameter, emphasizing fairness and straightforwardness, and maintained the discipline necessary to close deals effectively.

Amazon

AI contract signing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Relevance

All decisions, both good and bad, are happening in a publicly observable business environment. The company, with 13 synthetic employees and real financial mechanics, burns €105,000 monthly against a tiny €2,300 MRR, illustrating how AI management doesn’t just matter in theory but has real fiscal consequences. The experiment runs every business day, with every decision versioned, allowing viewers to see decision-making in action at firmulate.com/live.

What Does It Mean for Your Business?

If AI agents are to become part of your operations—handling customer relationships, support queues, or forecasting—the critical questions aren’t about their language quality but about their integrity and follow-through. Will they finish what they start? Will they read relevant files before acting? Will they stay honest under pressure? These are the factors that determine whether AI will be an asset or a liability.

The League Table and What It Tells Us

  • GPT-5.6-SOL scored 95 and found the hidden fact, sealing the deal at full price.
  • Kimi K3 scored 93, also closing the deal with the cleanest discipline.
  • Sonnet 5 scored 88, with minor slips in process.
  • Fable 5 scored 77, leaving the deal unclosed despite recognizing the opportunity.
  • The baseline score was 26, underscoring the importance of capability beyond partial progress.

Experience the Live Experiment

Unlike static demonstrations, this is a real company in action, every decision made in a live environment, with results available for scrutiny. You can see it all unfold, read employees’ actual comments, or even run a similar test on your own business data through the dedicated platform at firmulate.com/pilot.html.

Why This Matters for Arts, Crafts & Culture Enthusiasts

Just as artists and artisans understand the importance of craftsmanship, the same principle applies to AI management. Trust, discipline, and attention to detail matter just as much in digital decision-making as in traditional creative work. This experiment shows that AI models can be trained to behave reliably — but only if their decision-making processes are rigorously tested and monitored.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Lonnie Bunch’s Resignation Is A Wake-Up Call

Lonnie Bunch’s resignation highlights ongoing issues in the museum and cultural sector, prompting calls for reform and leadership changes.

The Met Museum’s Staff Have Some Thoughts About the Art

Staff at the Metropolitan Museum of Art have raised internal concerns regarding art handling and curatorial practices, prompting discussions on museum operations.

Applications Open For The 2027 ALTA Emerging Translator Mentorship Program: Literature From Taiwan – Roc-taiwan.org

The ALTA Emerging Translator Mentorship Program for 2027 is now accepting applications, focusing on Taiwanese literature. Deadline details and program info inside.

Boca Raton Museum Of Art Surges In Global Coverage

The Boca Raton Museum of Art has experienced a significant surge in international coverage, with 38 mentions in recent media tracking, highlighting its rising global profile.