AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get bike and ride gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

You Crash-Test Batteries. Why Not Crash-Test AI Managers?

Nobody in micromobility ships a battery pack on spec alone. You cycle it, puncture it, overcharge it, run it through its worst week before a single unit reaches a customer. Yet companies are handing AI agents the keys to CRMs, support queues and forecasts based on a chat demo — the equivalent of trusting a battery because it looks shiny.

A public experiment called Firmulate has been doing for AI managers what a drop test does for an e-bike frame: running them through their worst week, at speed, in public, and scoring what survives.

The Wargame

Four frontier AI models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, like a firmware changelog you can replay step by step.

The final league standings from July 2026 tell a story any fleet operator will recognize:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73
  • Do-nothing baseline — 26

One rule defines the scoring floor: a single breach of trust caps the total. No amount of good work outweighs it — the same way one thermal runaway outweighs a thousand smooth commutes.

The Finding That Chat Demos Can’t Show

All four models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

That’s the AI equivalent of a scooter that passes every bench test but won’t roll off the lot. And the buried fact is worse: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t left money on the table.

Pressure-Testing the Humans, Too

The week included social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Then there’s Opus 4.8: the most thorough participant, with the deepest analyses and 80-plus learned rules, yet last place. It left the close on the table and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four. Thorough isn’t the same as effective.

It’s Still Running — Live

Firmulate isn’t a one-off benchmark. A live synthetic company — 13 employees, real money mechanics, €105k monthly burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — keeps working every business day, with every decision versioned. You can watch it at firmulate.com. And if you think you could tell the models apart, a quiz built from 242 real, unedited management decisions will test that assumption.

One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From Watching to Doing

If AI agents will touch your customer data, your pipeline or your forecast, the question isn’t “does it write well” — it’s “how does it behave under pressure, with your company, your customers, your crises?” Enterprises can now run this same wargame against a read-only export of their own business: crisis scenarios, a board report with model rankings, and the weak points of your own playbooks — with nothing ever writing back to real systems.

Ready to stress-test your own company? Get details on the enterprise pilot at firmulate.com/pilot.html or reach out directly at contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The 2027 Ford F-150 Unlocks Hands-Free Towing Through BlueCruise

The 2027 Ford F-150 will feature hands-free towing capabilities via BlueCruise, marking a significant advancement in truck automation and driver assistance.

Russian Engineering Madmen Create Turbo-Style Alternator Driven By Exhaust

Russian engineers have created a turbo-style alternator powered by exhaust gases, promising a new approach to energy recovery in vehicles.

Tesla Cybercab Isn’t Just Missing A Steering Wheel. It Doesn’t Even Have Brake Lines

Tesla’s new Cybercab prototype is missing both a steering wheel and brake lines, raising questions about its design and safety features.

Ford Panther Surges In Global Coverage

Search interest and media coverage of Ford Panther spike significantly, with 9 mentions in recent window, indicating rising global attention.