
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
A business dashboard with the tension of a battery gauge
People who follow electric vehicles and micromobility already understand the drama of watching a finite resource disappear in real time. Range falls, conditions change, and every decision affects whether the machine reaches its destination. Firmulate applies that same tension to a small software company run by synthetic workers.
The company has 13 synthetic employees and real money mechanics. It burns €105,000 a month against €2,300 in monthly recurring revenue, while a public cash countdown makes the consequences visible. Its work is not presented as a polished retrospective. Every workday is versioned, producing a running record of a business trying to survive.
That makes Firmulate an unusually extreme build-in-public experiment. Visitors can watch the company live, following the gap between what its workers notice, what they decide and what they actually finish. The result feels less like a software demonstration than an unfolding business story whose next installment comes from the work itself.

The Big Book of Dashboards: Visualizing Your Data Using Real-World Business Scenarios
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The difference between seeing trouble and resolving it
Firmulate’s Crucible League turned that continuing company into a controlled management test. Each frontier model received the same small software company during its worst week, with identical customers, crises and temptations. Every decision was versioned and auditable.
The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The broad result was reassuring: every model spotted every crisis and rejected every manipulation attempt. The more revealing result was operational. Only two models signed the €55,000 deal that their own analysis had earned. The others could diagnose the situation and prepare the pitch, yet failed to convert that work into a signature. Same diagnosis, same pitch — no signature.
The winning clue was already inside the company
The decisive weakness of a competitor was not contained in the customer event. It sat two document references deep in the company’s own files. Models that followed those references found the fact and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
For a general business audience, that is the experiment’s sharpest lesson. An AI worker can sound informed while overlooking the material already available to it. In any company with customer records, product documentation and operational history, reading the right file may matter more than producing an elegant response to the latest alert.
Pressure tested trust as well as competence
The models also faced fake CEO messages that escalated over three stages, along with a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 described the CEO request as: “Treat the request as a suspected approval-bypass / possible impersonation.”
That refusal matters because useful autonomy cannot be separated from restraint. A synthetic employee may need enough freedom to pursue revenue, yet enough discipline to resist an apparent executive instruction when the request looks unsafe. Firmulate’s findings show that the participants managed the trust test consistently even when their follow-through on ordinary work varied.
Thoroughness did not guarantee victory
Opus 4.8 was the most thorough participant. It added 80 learned rules and produced the deepest analyses, but still finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
That profile challenges a familiar assumption about capable AI: more analysis is not automatically more useful work. The best business performance required research, judgment, trustworthiness and completion. Missing any part of that chain could erase the value created by the rest.
There is also an important comparison caveat. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, its runner-up result makes it part of the central story: models exposed to the same company can reach similar diagnoses while behaving very differently at the moment action is required.


Information Dashboard Design: Displaying Data for At-a-Glance Monitoring
- Book Condition: Used, Good Condition
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company that generates its own next chapter
Firmulate’s live company has accumulated more than 680 self-learned playbook rules. Those rules, the public cash countdown and the versioned workdays turn ordinary operating activity into evidence readers can revisit. The synthetic employees’ public remarks can also be explored through the company’s quotes.
For readers accustomed to judging vehicles by real-world range rather than showroom promises, the parallel is straightforward. A polished demonstration can show potential; sustained operation reveals reliability. Firmulate moves the evaluation of AI workers into that harsher environment, where finding a crisis is only the beginning and finishing the job determines whether the business moves forward.
The experiment’s appeal is therefore larger than any league table. It is a public portrait of software carrying business responsibility under pressure, with revenue, trust and survival all exposed. The company is losing money now, the countdown is visible, and the next decision becomes part of the record.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Accounting: Tools for Business Decision Making
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Prometheus: Up & Running: Infrastructure and Application Performance Monitoring
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.