Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Imagine a world where artificial intelligence isn’t just about chatbots and quick responses, but about reliably closing deals, reading critical files, and staying honest under pressure. For arts and culture organizations navigating complex decisions and delicate negotiations, understanding what truly separates an AI that can perform from one that only appears capable is essential. This is exactly what the latest experiment by Firmulate demonstrates—using AI to run a real company through its toughest week and revealing the unseen qualities that matter most.

Testing AI in the Real World of Business Challenges

In a groundbreaking live experiment, four of the world’s leading AI models were tasked with managing a small software company facing its worst week—full of crises, temptations, and critical decisions. The models, each with different capabilities and training, were given the same environment: the same customers, the same problems, and the same opportunities to manipulate or cheat.

What makes this test compelling is that it focuses on management skills—things like reading and interpreting vital company documents, resisting deceptive tactics, and following through on decisions—rather than just generating convincing chat responses. Every decision made by these models was documented, versioned, and auditable, ensuring transparency in their performance.

AI Change Management Made Simple: A 9-Step Framework for Business Leaders to Drive Generative AI Transformation (Reduce AI Fear, Win Buy-in, and Accelerate AI Adoption Across Your Organization)

AI Change Management Made Simple: A 9-Step Framework for Business Leaders to Drive Generative AI Transformation (Reduce AI Fear, Win Buy-in, and Accelerate AI Adoption Across Your Organization)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: The Surface vs. the Hidden

All four models successfully identified every crisis and refused every attempt at manipulation, including social engineering schemes like fake CEO messages and reporter tricks. This shows that current AI models are adept at recognizing problems and resisting deception—an essential skill in business and arts management.

However, the real difference lay beneath the surface. Only two models managed to close the deal that their own analysis had earned—the €55,000 contract—by executing the recommendations based on their own diagnosis. The other two, despite identifying the opportunities, left the deal on the table or failed to follow through.

This underscores a crucial insight: chat demos can hide a model’s true management capability. An AI’s ability to recognize a deal and actually close it—especially when it requires reading and acting on internal documents—is far more telling than what it can produce in a conversational snippet.

Free Fling File Transfer Software for Windows [PC Download]

Free Fling File Transfer Software for Windows [PC Download]

  • User-Friendly FTP Interface: Intuitive and easy to navigate
  • Reliable Site Maintenance: Ensures stable FTP connections
  • FTP Automation & Sync: Automates and synchronizes transfers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Decisive Weakness: Reading and Acting on Files

Further analysis revealed that the models which succeeded at closing the deal had read and understood documents buried two references deep within the company’s files—something that did not show up in their chat responses. The models that read these files effectively gained a comprehensive understanding of the context, enabling decisive action.

In contrast, models that didn’t delve into the company’s internal documents failed to recognize critical clues or to act decisively. This suggests that the true test of an AI’s management skill isn’t just in its conversational fluency but in its ability to read, interpret, and act on complex internal information.

Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support

Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Discipline Under Pressure: The Human-Like Test of Integrity

Another critical aspect was resistance to social engineering. All models refused staged attempts to get them to approve a fake CEO message or a reporter’s sneaky request—an encouraging sign that AI can be built to stay honest when under pressure.

For instance, Kimi K3 responded: “Treat the request as a suspected approval-bypass / possible impersonation,” demonstrating a cautious and disciplined approach. This kind of integrity is essential for AI systems that will handle sensitive business data or operate in environments where trust is paramount.

Amazon

AI internal document interpretation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Measure: Closing Strength, Not Just Chat Quality

What does this mean for arts, crafts, and cultural organizations considering AI tools? The lesson is clear: evaluating an AI’s capability based solely on chat quality is misleading. The real test is whether the AI can complete the work, read your files carefully, stay honest under pressure, and ultimately close the deals or make the right decisions.

The latest league table from the experiment ranks models by overall performance. The top performer, gpt-5.6-sol, achieved a score of 95, found the buried fact, and closed the deal—delivering full performance. Kimi K3 was close behind with a 93 score, also closing the deal with the cleanest discipline. The other models showed promise but faltered in execution, leaving opportunities unexploited or deals unclosed.

Implications for Arts and Culture Leaders

In arts and cultural sectors, where trust, authenticity, and delicate negotiations are everyday realities, this experiment underscores a vital point: not all AI models are equally capable of executing complex, trust-based tasks. An AI that reads your internal files, resists manipulation, and follows through on commitments can be a true partner—not just a chat engine.

Organizations should look beyond surface-level chat demos and consider whether their AI tools can perform these deeper management tasks. Running such real-world tests—the “wargame” of your own business—can reveal the true potential of AI in supporting your mission and operations.

Explore and Prepare with Firmulate

Want to see how your organization’s AI readiness stacks up? The Firmulate platform offers live experiments, tests, and pilots that simulate real business crises—fully transparent, auditable, and tailored to your needs. Discover whether your future AI workforce can read, decide, and deliver when it counts.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The vintage beauty of Soviet control rooms (2018)

A detailed look at Soviet-era control rooms, their design, technology, and cultural significance, based on the 2018 exploration of these historic sites.

Brett Whiteley Surges In Global Coverage

Brett Whiteley’s work is experiencing a surge in international coverage, with GDELT reporting 66 mentions in recent days, tripling baseline levels.

Show HN: Opening Lines Of Famous Literary Works

A developer shares a project collecting and displaying opening lines from renowned books to showcase literary beginnings daily.

Saman And Sasan Oskouei’s Sculptures Puncture The Weight Of Everyday Objects

Iranian artists Saman and Sasan Oskouei create sculptures that transform ordinary objects into thought-provoking art, puncturing perceptions of daily life.