
In the world of AI, performance metrics often seem like a black box of scores and percentages. But what if the baseline—doing nothing—still earns a significant score? That’s the surprising insight from a recent public benchmark that tests AI models in a realistic business setting. This experiment uncovers not just how AI models perform, but how their integrity and discipline matter just as much as their intelligence.
Get everyday essentials delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
The Real-World Test for AI in Business
Imagine a small software company faced with its worst week—similar customers, crises, and the temptations to cut corners. Now, picture different AI models managing this scenario, each with the same goal: navigate crises, avoid manipulation, and close deals. This isn’t a hypothetical; it’s a live experiment run by Firmulate, a leading platform that simulates real business environments for AI evaluation.
Every model is subjected to the same conditions: fake CEO messages, customer crises, and internal document references. The models’ decisions are transparent, auditable, and carefully analyzed. The goal? Measure management quality, not just chat skills or superficial performance.
AI ethics and integrity training courses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Baseline: 26 Points
In any evaluation, you might expect that doing nothing—ignoring crises and not engaging—would result in a score of zero. But the benchmark results defy this expectation. The so-called “do-nothing baseline” scores 26 points. Why? Because simply not engaging in manipulation attempts or crisis escalation counts as partial progress. It’s a recognition that honesty and discipline are foundational qualities, even if the AI doesn’t actively solve every problem.
Moreover, the scoring system caps the influence of breaches of trust. If an AI model acts unethically, it doesn’t just lose points; it hits a ceiling that prevents it from scoring higher—even if it performs well otherwise. This design ensures that integrity matters as much as raw capability, reflecting a realistic standard for AI in sensitive roles.
As an affiliate, we earn on qualifying purchases.
Beyond Surface-Level Performance
The experiment’s key finding is that while all models recognized crises and refused manipulative tactics, critical weaknesses determined ultimate success. The models that read deeper into internal documents—beyond just customer interactions—had the edge in closing deals at full price. The best performer, GPT-5.6, identified a buried fact in internal files and ultimately signed the deal worth over €4,580 monthly recurring revenue (MRR).
Surprisingly, even the most thorough model, Opus 4.8, finished last because it showed discipline lapses—writing attempts got locked away instead of escalating, and some opportunities were left on the table. The lesson? Deep analysis alone isn’t enough; consistency, discipline, and trustworthiness are vital.
As an affiliate, we earn on qualifying purchases.
Trust and Ethical Decision-Making Under Pressure
The benchmark also tested models against social engineering tactics—fake CEO messages escalating over several stages and a reporter trick requiring just a yes/no reply. All models refused to manipulate or impersonate, demonstrating a shared ethical stance. Kimi K3’s on-record reasoning encapsulates this: “Treat the request as a suspected approval-bypass / possible impersonation.”
This indicates that models can be programmed to recognize and resist ethically questionable prompts, an essential trait for AI operating in business environments.
As an affiliate, we earn on qualifying purchases.
Implications for Business AI Adoption
The experiment underscores a crucial point: it’s not enough for AI to generate convincing text or solve problems superficially. What truly matters is whether these models can finish what they start, stay honest under pressure, and read internal files that contain critical information. These qualities determine if an AI can be trusted to handle real-world business tasks without risking reputation or compliance breaches.
While the scores seem modest—the top score being 95—the key takeaway is the importance of integrity. A model that cheats or slips in discipline may appear competent but fails at the core of trustworthy management. That’s why the benchmark includes a “floor” score of 26, representing a baseline of honesty that any AI should surpass to be viable for serious deployment.
Watch the Experiment Live
Interested in seeing this in action? The live setup at firmulate.com/live allows you to observe how different AI models handle real crises, decisions, and manipulations—replicating a genuine business environment. Every decision and outcome is visible, providing transparency and a clear picture of what trustworthy AI management looks like in practice.
In an era where AI’s role in business is expanding rapidly, understanding these benchmarks equips decision-makers to choose models that are not only capable but also disciplined and honest. The message from this experiment is clear: in AI, integrity isn’t optional. It’s fundamental.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
