
Get bike and ride gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
You Crash-Test Batteries. Why Not Crash-Test AI Managers?
Nobody in micromobility ships a battery pack on spec alone. You cycle it, puncture it, overcharge it, run it through its worst week before a single unit reaches a customer. Yet companies are handing AI agents the keys to CRMs, support queues and forecasts based on a chat demo — the equivalent of trusting a battery because it looks shiny.
A public experiment called Firmulate has been doing for AI managers what a drop test does for an e-bike frame: running them through their worst week, at speed, in public, and scoring what survives.
The Wargame
Four frontier AI models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, like a firmware changelog you can replay step by step.
The final league standings from July 2026 tell a story any fleet operator will recognize:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
- Do-nothing baseline — 26
One rule defines the scoring floor: a single breach of trust caps the total. No amount of good work outweighs it — the same way one thermal runaway outweighs a thousand smooth commutes.
The Finding That Chat Demos Can’t Show
All four models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
That’s the AI equivalent of a scooter that passes every bench test but won’t roll off the lot. And the buried fact is worse: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t left money on the table.
Pressure-Testing the Humans, Too
The week included social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then there’s Opus 4.8: the most thorough participant, with the deepest analyses and 80-plus learned rules, yet last place. It left the close on the table and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four. Thorough isn’t the same as effective.
It’s Still Running — Live
Firmulate isn’t a one-off benchmark. A live synthetic company — 13 employees, real money mechanics, €105k monthly burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — keeps working every business day, with every decision versioned. You can watch it at firmulate.com. And if you think you could tell the models apart, a quiz built from 242 real, unedited management decisions will test that assumption.
One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second.

From Watching to Doing
If AI agents will touch your customer data, your pipeline or your forecast, the question isn’t “does it write well” — it’s “how does it behave under pressure, with your company, your customers, your crises?” Enterprises can now run this same wargame against a read-only export of their own business: crisis scenarios, a board report with model rankings, and the weak points of your own playbooks — with nothing ever writing back to real systems.
Ready to stress-test your own company? Get details on the enterprise pilot at firmulate.com/pilot.html or reach out directly at contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
