
Imagine trusting an AI to run your art gallery, handle your rare collectibles, or manage a delicate craft workshop. Would you accept an AI that scores just 26 out of 100 on honesty and reliability? Surprisingly, in the world of AI benchmarking, that’s precisely what a do-nothing baseline can achieve — revealing why trustworthiness isn’t just a nice-to-have but a quantifiable standard.
Get art and craft supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Reality of AI Performance in Business Settings
At first glance, one might expect a “do-nothing” AI baseline—an agent that simply stays idle—to score zero, indicating no competence or effort. But in the recent Firmulate experiment, this baseline surprisingly scored 26 out of 100. Why? Because even minimal operations, such as reading files or recognizing crises, count as partial progress. This baseline underscores a critical point: AI performance isn’t just about action, but about integrity and the capacity to avoid pitfalls.
AI ethics and trustworthiness software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Understanding the Benchmark Methodology
In the live experiment, each AI model was tasked with managing a small software company facing its worst week—same customers, crises, and temptations. Every decision was carefully versioned and auditable, ensuring transparency. What’s revealing is how different models handled similar scenarios:
- All models identified every crisis and refused manipulation attempts, demonstrating strong integrity.
- Only two models managed to close the actual deal, earning €55,000 in revenue—their own analysis led them to the right conclusion, and they signed the contract.
- Interestingly, the key weakness wasn’t in customer interactions but in internal document reading. Those models that examined files deeply uncovered opportunities that others missed, securing additional revenue potential of over €4,500 monthly recurring revenue.
AI document reading and analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Trust Matters More Than Just Being Correct
One of the most striking parts of this benchmark is how models responded to social engineering attempts, such as fake CEO messages or reporter tricks. All five models refused to escalate or approve suspicious requests, citing suspicion of impersonation or approval bypass. This demonstrates that AI’s ability to recognize manipulation and act ethically is critical—especially when managing sensitive or high-stakes operations.
AI manipulation detection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human-Like Challenges and AI Discipline
The live environment involved 13 synthetic employees operating under real money mechanics—burning €105,000 monthly against a revenue of €2,300. The system uses over 680 self-learned rules, with every decision versioned for accountability. Even the most thorough participant, Opus 4.8, scored the lowest—highlighting that thoroughness alone isn’t enough. Discipline slips, like leaving decisions unescalated or failing to close opportunities, can cost real revenue and trust.
As an affiliate, we earn on qualifying purchases.
What the Numbers Reveal for Arts and Culture
While the experiment centers on business operations, the lessons extend to arts, crafts, and cultural management. If AI is to assist in curating exhibitions, managing rare collections, or engaging with audiences, it must do more than just produce appealing content. It must read and interpret internal documents, recognize manipulative behavior, and uphold trustworthiness—even under pressure. A benchmark score of 26 for a do-nothing baseline illustrates that minimal effort isn’t enough: AI must be aligned with human values and ethical standards to be truly useful.
The Takeaway: Trust and Performance Go Hand in Hand
The key insight from the Firmulate benchmark is that honesty and discipline are fundamental metrics—more telling than superficial chat quality or cleverness. The fact that a baseline can score 26 points, simply by reading files and refusing manipulation, highlights how essential it is to develop AI that can read, interpret, and act ethically at a foundational level. As AI begins to touch more aspects of art management and cultural curation, these standards will serve as a guide for choosing trustworthy, reliable tools.
Try It Yourself and See the Future of AI in Your Field
Interested in exploring how your organization might fare? Firms can run the same scenario against their own data, with no risk to actual systems. This pilot approach allows businesses to assess their AI’s decision-making quality before full deployment. Visit firmulate.com/pilot.html to learn more about this risk-free testing method.

In AI management, trustworthiness isn’t just about accuracy—it’s about integrity, reading internal documents, and resisting manipulation. The Firmulate benchmark proves even a do-nothing baseline scores 26, emphasizing the importance of discipline and ethical behavior in AI systems.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
