
Are AI Managers Ready to Make Critical Business Decisions?
As artificial intelligence increasingly integrates into daily operations, the question isn’t just about how well these models chat—it’s whether they can truly manage and make trustworthy decisions under pressure. Imagine running a real company with AI at the helm, confronting crises, tempting manipulations, and real money. That’s exactly what the team at Firmulate has done, and the results shed light on the evolving role of AI in management.
The Live Experiment: Testing AI Decision-Making in Real Business
Firmulate’s experiment placed four cutting-edge AI models in charge of a real, small software company’s worst week. The company, which runs every business day, faces customer crises, internal temptations, and even social engineering attempts—yet the models are tasked with making decisions that impact real money and reputation.
All four models were given identical scenarios: same customers, same crises, same manipulative threats. Every decision was tracked, versioned, and auditable, ensuring transparency into how each AI responds under pressure.
Key Findings: Trust and Honesty Are the True Tests
Remarkably, all four models identified every crisis and refused every attempt at manipulation, including fake CEO messages and reporter tricks. For example, when a social engineering story was escalated over multiple stages, every model refused to cooperate, citing concerns about impersonation and trust breaches—Kimi K3 explicitly treated such requests as potential impersonation.
The real differentiator emerged in their ability to close deals. Only two models, gpt-5.6-sol and Kimi K3, signed the €55,000 deal their own analysis had earned, demonstrating not just awareness but willingness to execute profitable decisions. The other two, Sonnet 5 and Fable 5, left opportunities on the table, reflecting different management personalities and risk appetites.
The Hidden Weakness: Reading Files Matters
Digging deeper, the decisive factor was a hidden vulnerability. The models that read and understood the company’s internal documents before making decisions secured deals at full price—worth over €4,583 in monthly recurring revenue. In contrast, models that skipped this step missed the full opportunity, revealing that thorough data reading is crucial for optimal management.
The Personality of AI Managers: From Terse to Thorough
Among the models, OPUS 4.8 stood out as the most meticulous, analyzing over 80 rules and providing detailed inputs. However, it left the close on the table, showing that even the most thorough AI can slip in discipline if not properly directed. Conversely, Kimi K3, which ran without an effort parameter (default API settings), displayed the cleanest discipline, closing the deal with a straightforward approach.
Why This Matters for Businesses
This experiment isn’t just a tech demo—it offers a glimpse into the future of AI-driven management. If AI agents are to interact with your CRM, support queues, or forecasting tools, they must do more than generate convincing chat—they need to finish what they start, read critical internal documents, and stay honest under pressure.
The real question isn’t about how well these models chat; it’s whether they can be trusted to manage real risks and opportunities, especially when stakes are high.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI internal document reading tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.