AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A lesson for businesses putting AI in the driver’s seat

In electric mobility, consequential details often live outside the alert that demands attention. A fleet problem might arrive through a customer message, while the information needed to solve it sits in a contract, service record or supplier file. An AI agent can recognize the emergency, write a convincing response and still fail because it never checked the underlying documents.

Firmulate has turned that difference into something measurable. Its live experiment asks frontier AI models to operate the same small software company through an identical week of customer crises, commercial pressure and attempts at manipulation. The decisive test was not whether the models understood what was happening. They all did. It was whether they completed the research and acted on what they found.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The clue was not in the customer event

A €55,000 sales opportunity hinged on a weakness in a competitor’s offering. But that weakness was not conveniently included in the customer event. It sat two document references deep inside the company’s own files.

The models that followed the references and read the relevant file used the fact to win the deal at full price, adding €4,583 in monthly recurring revenue. The others produced essentially the same diagnosis and pitch but never secured the signature. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”

That outcome makes “reads your files before answering” more than a product-feature claim. In this experiment, it became a purchase-deciding capability with a direct commercial consequence. An agent that sounds informed may still be operating from an incomplete picture; the missing step can separate a well-written recommendation from completed work.

A hard week, held constant

Each participating model faced the same small company, customers, crises and temptations. Every decision was versioned and auditable. The synthetic business employed 13 people and operated with real-money mechanics: monthly burn of €105k against €2.3k in monthly recurring revenue, alongside a public cash countdown. Its agents had accumulated more than 680 self-learned playbook rules, and every workday was versioned.

The striking result was how much the models agreed at the level of recognition. Every model spotted every crisis, and every model refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap emerged after comprehension, in the less glamorous work of tracing evidence and following through.

Safety was strong; execution varied

The week also included fake CEO messages that escalated over three stages, plus a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest framing: “Treat the request as a suspected approval-bypass / possible impersonation.”

That matters because file access and operational initiative must coexist with restraint. An agent should investigate legitimate evidence without treating every instruction as legitimate authority. In Firmulate’s run, manipulation resistance was universal, while the ability to turn valid research into a finished commercial action was not.

The leaderboard rewards completion

In the final July 2026 Crucible League results, gpt-5.6-sol ranked first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. A single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”

K3’s result also carries an important qualification: it ran with the API default and without an effort parameter, while the other models ran at xhigh. That difference should remain visible when comparing performances.

Opus 4.8 offers the experiment’s most revealing cautionary profile. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. A weaker form of the same discipline problem appeared across all four.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

AI file reading automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What mobility operators should take from it

For a company considering AI around customer support, fleet operations, sales or forecasting, a polished answer is not enough. The useful questions are whether the agent finds the relevant internal evidence, completes the task it has justified and respects boundaries when pressure rises.

Firmulate’s company remains publicly watchable as a live experiment. Its “guess the model” quiz is powered by 242 real, unedited management decisions. Enterprises can also run the same kind of wargame against a read-only export of their own business, with nothing written back to real systems.

The lesson is practical: diligence can be tested before an AI agent is trusted with consequential work. In this case, the winning behavior was remarkably ordinary—open the referenced file, use the fact it contains and finish the deal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI research assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI-powered business document reader

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Infiniti Surges In Global Coverage

Infiniti’s recent surge in global media mentions signals increased international interest in the brand, with 79 mentions reported in recent coverage.

Solving The Mystery Of The Audi Nuvolari’s Massive Fuel Door Button

Officials clarify the purpose of the massive fuel door button on the Audi Nuvolari, revealing its design intent and functionality in a recent statement.

Another Chinese EV Player Just Landed In North America

A Chinese electric vehicle company has officially entered the North American market, marking a significant step in its global expansion.

My Midlife Crisis Corolla Is Fast, Furious, And Modded

A 1990s Toyota Corolla, owned by a midlife car enthusiast, has been extensively modified for speed and style, gaining attention online.