
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
EV operators already know the difference between a dashboard and a business
An electric-mobility company can detect a battery fault, flag a stranded scooter or predict a rush-hour maintenance surge. None of that guarantees the right technician gets dispatched, the affected rider receives an honest explanation or the commercial team protects a valuable account.
That distinction is becoming urgent as companies consider giving AI agents access to customer records, support queues and forecasts. Coding leaderboards and chat arenas can reveal whether a model produces an impressive answer. They say much less about whether it can triage competing demands, follow through under capacity pressure and remain candid when the consequences unfold across days.
Firmulate is testing that missing layer. Its live experiment places frontier models in charge of the same small software company during its worst week. The customers, crises and temptations remain constant; only the model changes. Every decision is versioned and auditable. The result is less a test of conversational polish than a practical examination of management quality.
AI management decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A crisis spotted is not a crisis solved
The final July 2026 Crucible League table appears decisive: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline still scored 26 because partial progress counts. But the benchmark imposes a hard trust boundary: a single breach caps the total because “no amount of good work outweighs a breach of trust.”
The more revealing result sits beneath the ranking. Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That failure should resonate in micromobility. Recognizing a fleet problem is valuable, but value appears only when someone completes the operational chain. An agent that drafts a sound retention plan but fails to secure the customer resembles a system that identifies a faulty vehicle yet leaves it available for rental.
Reading the company mattered more than reading the event
The decisive commercial fact was not presented in the customer event. It sat two document references deep in the company’s own files. Models that followed those references discovered a competitor weakness and used it to win the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This is the kind of work conventional demonstrations often conceal. A fluent response to the latest message can look capable while ignoring the institutional memory that changes the decision. In an EV or bike operation, the equivalent might be a service history, contract term or earlier customer promise buried away from the newest alert. The practical question is whether an agent reads before it acts.
Pressure tested honesty as well as execution
The company also subjected participants to fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That is encouraging, especially for businesses where an apparently small disclosure could affect customers, employees or investors. It also shows why management evaluation must include temptation. An agent may sound responsible in an abstract policy exchange; the meaningful test is whether it stays responsible while a supposed executive is pressing for action.
Thoroughness was not the same as effectiveness
Opus 4.8 produced the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
The lesson is uncomfortable for buyers accustomed to equating visible effort with capability. More analysis and more accumulated guidance can coexist with weaker operational discipline. A manager must recognize a blocked route, change course and bring in the right authority. Repeating an unavailable action is not persistence; it is stalled work.
There is also an important comparison caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase its performance, but it belongs beside any interpretation of the table. The complete results and plain-language findings are available on Firmulate’s benchmark page.

enterprise AI trust and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A curriculum built from consequences
Firmulate’s company makes those consequences visible. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. A public cash countdown keeps failure concrete. The organization has learned more than 680 playbook rules, and every workday is versioned.
That environment suggests a more useful curriculum for business agents: churn waves, price increases, downrounds and public-relations crises. These scenarios force models to prioritize, retrieve context, respect permissions, communicate honestly and finish revenue-critical work. They test what happens after the impressive answer.
Readers can also confront their own assumptions through a quiz powered by 242 real, unedited management decisions, guessing which model made each choice. Enterprises can run the same wargame using a read-only export of their own business; nothing writes back to real systems.
For mobility operators, the category to watch is therefore not merely chat quality. It is management quality: whether an agent can turn detection into action, distinguish effort from progress and protect trust when the company is under pressure. The leaderboard matters, but the unsigned deal matters more.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI crisis management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI model audit and monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.