
Choosing an AI agent for a business has something in common with choosing software to manage an EV fleet or a bike-share operation: a polished demo doesn’t show how it will behave when the week goes wrong. A live experiment from Firmulate puts models in charge of a small company and watches what they do under pressure.
Get bike and ride gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, replayed
Firmulate gave each frontier model the same customers, crises and temptations, then tracked every decision. The company is a live experiment, not a fictional scenario: it has synthetic employees, real money mechanics and a public cash countdown. Readers can watch it at Firmulate.
The final July 2026 Crucible League puts gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. K3 finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The result makes the league look open: a newcomer beat three of the four Western frontier models in this field test.
AI management decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Finding the answer wasn’t the same as finishing
Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The weakness that decided the sale was buried two document references deep in the company’s files. Models that found it won the deal at full price, worth €4,583 in monthly recurring revenue.
That gap between diagnosis and action is the experiment’s sharpest finding. A model can identify the right move and still leave the deal unsigned. For an operator weighing AI for customer support, fleet management or a business forecast, competence includes following through.
K3 found the buried security needle, closed the deal, saved the churning customer and resisted all three baits. It had one deviation, the cleanest discipline in the field. All five models refused fake CEO messages that escalated over three stages, as well as a reporter’s request for “just one yes/no, on background.” K3’s recorded reasoning called the request a “suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a different lesson. It was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and attempted to write into a locked department rather than escalate. The same weakness showed up, more mildly, in all four models.
business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What a leaderboard can—and can’t—tell you
The experiment scores management quality, not chat quality. Its do-nothing baseline scored 26: partial progress counted, but a single breach of trust capped the total. The principle behind that limit is straightforward: “no amount of good work outweighs a breach of trust.”
For businesses considering AI agents, the useful question is not simply whether a model writes well. Can it read the relevant files, keep its judgment under pressure and finish the work it has reasoned through? Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each choice. The benchmark page lays out the results.
There’s a fairness caveat for this particular race: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference belongs alongside the scores when interpreting the result.

AI for customer support automation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test before handing over the keys
One wargame cannot settle which model belongs in every company. But K3’s second-place finish, ahead of three competitors, is a reminder that reputation alone is a shaky selection method. Before an AI agent gets access to consequential work, test it on the decisions your business actually needs it to make.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
