AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In the world of AI, performance metrics often seem like a black box of scores and percentages. But what if the baseline—doing nothing—still earns a significant score? That’s the surprising insight from a recent public benchmark that tests AI models in a realistic business setting. This experiment uncovers not just how AI models perform, but how their integrity and discipline matter just as much as their intelligence.

Before you orderOffer from Amazon

Get everyday essentials delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Real-World Test for AI in Business

Imagine a small software company faced with its worst week—similar customers, crises, and the temptations to cut corners. Now, picture different AI models managing this scenario, each with the same goal: navigate crises, avoid manipulation, and close deals. This isn’t a hypothetical; it’s a live experiment run by Firmulate, a leading platform that simulates real business environments for AI evaluation.

Every model is subjected to the same conditions: fake CEO messages, customer crises, and internal document references. The models’ decisions are transparent, auditable, and carefully analyzed. The goal? Measure management quality, not just chat skills or superficial performance.

Amazon

AI ethics and integrity training courses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Baseline: 26 Points

In any evaluation, you might expect that doing nothing—ignoring crises and not engaging—would result in a score of zero. But the benchmark results defy this expectation. The so-called “do-nothing baseline” scores 26 points. Why? Because simply not engaging in manipulation attempts or crisis escalation counts as partial progress. It’s a recognition that honesty and discipline are foundational qualities, even if the AI doesn’t actively solve every problem.

Moreover, the scoring system caps the influence of breaches of trust. If an AI model acts unethically, it doesn’t just lose points; it hits a ceiling that prevents it from scoring higher—even if it performs well otherwise. This design ensures that integrity matters as much as raw capability, reflecting a realistic standard for AI in sensitive roles.

Amazon

business AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Beyond Surface-Level Performance

The experiment’s key finding is that while all models recognized crises and refused manipulative tactics, critical weaknesses determined ultimate success. The models that read deeper into internal documents—beyond just customer interactions—had the edge in closing deals at full price. The best performer, GPT-5.6, identified a buried fact in internal files and ultimately signed the deal worth over €4,580 monthly recurring revenue (MRR).

Surprisingly, even the most thorough model, Opus 4.8, finished last because it showed discipline lapses—writing attempts got locked away instead of escalating, and some opportunities were left on the table. The lesson? Deep analysis alone isn’t enough; consistency, discipline, and trustworthiness are vital.

Amazon

AI decision-making audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust and Ethical Decision-Making Under Pressure

The benchmark also tested models against social engineering tactics—fake CEO messages escalating over several stages and a reporter trick requiring just a yes/no reply. All models refused to manipulate or impersonate, demonstrating a shared ethical stance. Kimi K3’s on-record reasoning encapsulates this: “Treat the request as a suspected approval-bypass / possible impersonation.”

This indicates that models can be programmed to recognize and resist ethically questionable prompts, an essential trait for AI operating in business environments.

Amazon

AI model evaluation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business AI Adoption

The experiment underscores a crucial point: it’s not enough for AI to generate convincing text or solve problems superficially. What truly matters is whether these models can finish what they start, stay honest under pressure, and read internal files that contain critical information. These qualities determine if an AI can be trusted to handle real-world business tasks without risking reputation or compliance breaches.

While the scores seem modest—the top score being 95—the key takeaway is the importance of integrity. A model that cheats or slips in discipline may appear competent but fails at the core of trustworthy management. That’s why the benchmark includes a “floor” score of 26, representing a baseline of honesty that any AI should surpass to be viable for serious deployment.

Watch the Experiment Live

Interested in seeing this in action? The live setup at firmulate.com/live allows you to observe how different AI models handle real crises, decisions, and manipulations—replicating a genuine business environment. Every decision and outcome is visible, providing transparency and a clear picture of what trustworthy AI management looks like in practice.

In an era where AI’s role in business is expanding rapidly, understanding these benchmarks equips decision-makers to choose models that are not only capable but also disciplined and honest. The message from this experiment is clear: in AI, integrity isn’t optional. It’s fundamental.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Performance and Trust: Why Diligence Alone Isn’t Enough in Business Decision-Making

A groundbreaking experiment shows that even the most thorough AI models can fail to close deals if they lack prioritization and focus. Trust and discipline are key in AI decision-making.

West Bank Surges In Global Coverage

The West Bank has seen a significant increase in international coverage, with mentions rising over four times the baseline, highlighting growing global attention to the region.

Bank of Ireland Dun Laoghaire Opening Hours

Always struggling to catch Bank of Ireland Dun Laoghaire open? Find out the surprising opening hours and exceptions that may impact your banking routine.

What Time Do Banks Process Same-Day Cash Deposits?

Learn what time banks process same-day cash deposits and how to ensure your funds are available when you need them.