AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine trusting an AI to run your art gallery, handle your rare collectibles, or manage a delicate craft workshop. Would you accept an AI that scores just 26 out of 100 on honesty and reliability? Surprisingly, in the world of AI benchmarking, that’s precisely what a do-nothing baseline can achieve — revealing why trustworthiness isn’t just a nice-to-have but a quantifiable standard.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get art and craft supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Reality of AI Performance in Business Settings

At first glance, one might expect a “do-nothing” AI baseline—an agent that simply stays idle—to score zero, indicating no competence or effort. But in the recent Firmulate experiment, this baseline surprisingly scored 26 out of 100. Why? Because even minimal operations, such as reading files or recognizing crises, count as partial progress. This baseline underscores a critical point: AI performance isn’t just about action, but about integrity and the capacity to avoid pitfalls.

Amazon

AI ethics and trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark Methodology

In the live experiment, each AI model was tasked with managing a small software company facing its worst week—same customers, crises, and temptations. Every decision was carefully versioned and auditable, ensuring transparency. What’s revealing is how different models handled similar scenarios:

  • All models identified every crisis and refused manipulation attempts, demonstrating strong integrity.
  • Only two models managed to close the actual deal, earning €55,000 in revenue—their own analysis led them to the right conclusion, and they signed the contract.
  • Interestingly, the key weakness wasn’t in customer interactions but in internal document reading. Those models that examined files deeply uncovered opportunities that others missed, securing additional revenue potential of over €4,500 monthly recurring revenue.
Amazon

AI document reading and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Trust Matters More Than Just Being Correct

One of the most striking parts of this benchmark is how models responded to social engineering attempts, such as fake CEO messages or reporter tricks. All five models refused to escalate or approve suspicious requests, citing suspicion of impersonation or approval bypass. This demonstrates that AI’s ability to recognize manipulation and act ethically is critical—especially when managing sensitive or high-stakes operations.

Amazon

AI manipulation detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Human-Like Challenges and AI Discipline

The live environment involved 13 synthetic employees operating under real money mechanics—burning €105,000 monthly against a revenue of €2,300. The system uses over 680 self-learned rules, with every decision versioned for accountability. Even the most thorough participant, Opus 4.8, scored the lowest—highlighting that thoroughness alone isn’t enough. Discipline slips, like leaving decisions unescalated or failing to close opportunities, can cost real revenue and trust.

Amazon

AI decision auditing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Numbers Reveal for Arts and Culture

While the experiment centers on business operations, the lessons extend to arts, crafts, and cultural management. If AI is to assist in curating exhibitions, managing rare collections, or engaging with audiences, it must do more than just produce appealing content. It must read and interpret internal documents, recognize manipulative behavior, and uphold trustworthiness—even under pressure. A benchmark score of 26 for a do-nothing baseline illustrates that minimal effort isn’t enough: AI must be aligned with human values and ethical standards to be truly useful.

The Takeaway: Trust and Performance Go Hand in Hand

The key insight from the Firmulate benchmark is that honesty and discipline are fundamental metrics—more telling than superficial chat quality or cleverness. The fact that a baseline can score 26 points, simply by reading files and refusing manipulation, highlights how essential it is to develop AI that can read, interpret, and act ethically at a foundational level. As AI begins to touch more aspects of art management and cultural curation, these standards will serve as a guide for choosing trustworthy, reliable tools.

Try It Yourself and See the Future of AI in Your Field

Interested in exploring how your organization might fare? Firms can run the same scenario against their own data, with no risk to actual systems. This pilot approach allows businesses to assess their AI’s decision-making quality before full deployment. Visit firmulate.com/pilot.html to learn more about this risk-free testing method.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

In AI management, trustworthiness isn’t just about accuracy—it’s about integrity, reading internal documents, and resisting manipulation. The Firmulate benchmark proves even a do-nothing baseline scores 26, emphasizing the importance of discipline and ethical behavior in AI systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Nairobi Architecture City Guide: 11 Projects Building A Post-Independence Identity – ArchDaily

Exploring 11 significant architectural projects in Nairobi shaping its post-independence identity, reflecting cultural and urban development since independence.

AI Showdown Reveals the Hidden Skill That Separates the Good from the Great in Business Execution

A live experiment shows that AI’s real management skills—reading, resisting manipulation, and closing deals—are invisible in chat demos but vital for trust and success in arts and culture.

Arts Centre Melbourne Surges In Global Coverage

Arts Centre Melbourne has experienced a surge in international media coverage, with 28 mentions in recent reports, highlighting its growing global profile.

Hanoi Unveils Cultural And Architectural Heritage Of 54 Ethnic Groups

Hanoi launches a comprehensive review highlighting the cultural and architectural diversity of 54 ethnic groups, emphasizing preservation and tourism.