
A convincing answer is easy to publish. A reliable decision is harder to judge. For newsrooms and media businesses weighing AI agents, the sharper question is what happens when a model faces pressure, incomplete information and a decision with real consequences.
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate puts that question into a live experiment: several AI models run the same small software company through its worst week. The exercise offers a useful lens for media leaders considering agents for customer support, subscriptions or business operations.
One company, the same hard week
In the final Crucible League, published in July 2026, five participants were ranked: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. The experiment gave each frontier model the same customers, crises and temptations. Decisions were versioned and auditable.
The results point beyond whether a model can produce a polished response. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the finding put it: “Same diagnosis, same pitch — no signature.”
The important clue was buried
The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. For a media company, the parallel is familiar: the detail that changes a decision may live in a contract, a subscriber note or an internal policy, rather than in the latest message.
The trial also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s familiar-sounding request: “just one yes/no, on background”. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a telling standard for any organization where a plausible source or senior-sounding request can exert pressure.
Thoroughness did not guarantee the close
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four participants. Careful analysis and sound instincts did not consistently turn into a completed, controlled action.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The leaderboard is a snapshot of this experiment, not a universal verdict on which model is best for every newsroom or business.
A live company, and a route to a pilot
The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and a versioned record for every workday. The experiment is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.
For enterprise teams, the next step is a pilot against a read-only export of their own business. They can examine crisis scenarios, model rankings and weak points in their playbooks without writing back to real systems. For a newsroom, that could make the question concrete: how would an agent handle a subscription crisis, a sensitive customer request or a pressure campaign using the information the organization actually holds?

Move from watching to testing
A live benchmark can show how models behave in someone else’s company. A pilot can show what their decisions look like against yours. To explore a read-only business wargame, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
