AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

EV operators already know the difference between a dashboard and a business

An electric-mobility company can detect a battery fault, flag a stranded scooter or predict a rush-hour maintenance surge. None of that guarantees the right technician gets dispatched, the affected rider receives an honest explanation or the commercial team protects a valuable account.

That distinction is becoming urgent as companies consider giving AI agents access to customer records, support queues and forecasts. Coding leaderboards and chat arenas can reveal whether a model produces an impressive answer. They say much less about whether it can triage competing demands, follow through under capacity pressure and remain candid when the consequences unfold across days.

Firmulate is testing that missing layer. Its live experiment places frontier models in charge of the same small software company during its worst week. The customers, crises and temptations remain constant; only the model changes. Every decision is versioned and auditable. The result is less a test of conversational polish than a practical examination of management quality.

Amazon

AI management decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A crisis spotted is not a crisis solved

The final July 2026 Crucible League table appears decisive: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline still scored 26 because partial progress counts. But the benchmark imposes a hard trust boundary: a single breach caps the total because “no amount of good work outweighs a breach of trust.”

The more revealing result sits beneath the ranking. Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That failure should resonate in micromobility. Recognizing a fleet problem is valuable, but value appears only when someone completes the operational chain. An agent that drafts a sound retention plan but fails to secure the customer resembles a system that identifies a faulty vehicle yet leaves it available for rental.

Reading the company mattered more than reading the event

The decisive commercial fact was not presented in the customer event. It sat two document references deep in the company’s own files. Models that followed those references discovered a competitor weakness and used it to win the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is the kind of work conventional demonstrations often conceal. A fluent response to the latest message can look capable while ignoring the institutional memory that changes the decision. In an EV or bike operation, the equivalent might be a service history, contract term or earlier customer promise buried away from the newest alert. The practical question is whether an agent reads before it acts.

Pressure tested honesty as well as execution

The company also subjected participants to fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

That is encouraging, especially for businesses where an apparently small disclosure could affect customers, employees or investors. It also shows why management evaluation must include temptation. An agent may sound responsible in an abstract policy exchange; the meaningful test is whether it stays responsible while a supposed executive is pressing for action.

Thoroughness was not the same as effectiveness

Opus 4.8 produced the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

The lesson is uncomfortable for buyers accustomed to equating visible effort with capability. More analysis and more accumulated guidance can coexist with weaker operational discipline. A manager must recognize a blocked route, change course and bring in the right authority. Repeating an unavailable action is not persistence; it is stalled work.

There is also an important comparison caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase its performance, but it belongs beside any interpretation of the table. The complete results and plain-language findings are available on Firmulate’s benchmark page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI trust and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A curriculum built from consequences

Firmulate’s company makes those consequences visible. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. A public cash countdown keeps failure concrete. The organization has learned more than 680 playbook rules, and every workday is versioned.

That environment suggests a more useful curriculum for business agents: churn waves, price increases, downrounds and public-relations crises. These scenarios force models to prioritize, retrieve context, respect permissions, communicate honestly and finish revenue-critical work. They test what happens after the impressive answer.

Readers can also confront their own assumptions through a quiz powered by 242 real, unedited management decisions, guessing which model made each choice. Enterprises can run the same wargame using a read-only export of their own business; nothing writes back to real systems.

For mobility operators, the category to watch is therefore not merely chat quality. It is management quality: whether an agent can turn detection into action, distinguish effort from progress and protect trust when the company is under pressure. The leaderboard matters, but the unsigned deal matters more.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI crisis management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model audit and monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Jeff Gordon And Bill Elliott Raced IROC Pontiacs Wheel-to-Wheel In DC Last Weekend. Who Knew?

Jeff Gordon and Bill Elliott competed wheel-to-wheel in IROC Pontiacs last weekend in Washington, D.C., in a rare exhibition event that drew attention from racing fans.

This Mitsubishi Pajero Mini Is A Factory JDM Snoopy Edition

Mitsubishi has officially launched a limited-edition Pajero Mini featuring a Snoopy-themed design exclusively for the Japanese domestic market.

How Much Quicker Is The Newest Version Of Tesla FSD? This Test Put It Against Old Software To Find Out

Tesla’s latest FSD Beta V12 was tested against its previous version, showing significant improvements in response time and driving efficiency, according to recent tests.

Hyundai Surges In Global Coverage

Hyundai experiences a surge in international coverage, with 54 media mentions in recent analysis, highlighting increased global interest in the automaker.