
The Demo Isn’t the Job
Anyone who has spent time in a workshop knows the routine. The catalog says a saw has a 15-amp motor and a laser guide. Then you rip eight feet of maple, the fence drifts a sixteenth, and you learn what the brochure never told you: specs aren’t the same as performance under load.
Buyers of AI agents are learning the same lesson the hard way. The industry’s favorite measuring sticks — coding leaderboards and chat arenas — grade how well a model answers. They don’t grade whether it finishes what it starts, reads your files before acting, or stays honest when someone tries to con it. That’s the gap a live experiment at Firmulate set out to measure, and the results are worth a look even if you’ve never written a line of code.
The Worst Week in Business, Four Times Over
Firmulate ran four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 in the final league — through an identical gauntlet: each was put in charge of the same small software company during its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so there’s no arguing after the fact about who did what.
The final Crucible League standings from July 2026: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. A do-nothing baseline scored 26 — partial progress counts for something, but a single breach of trust caps the total. In Firmulate’s words, “no amount of good work outweighs a breach of trust.”
Same Diagnosis, Same Pitch — No Signature
Here’s the finding that chat demos can’t show. Every model spotted every crisis. Every model refused every manipulation attempt. But only two of them actually closed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
Imagine a sales rep who nails the diagnosis, delivers a flawless presentation, and then just… doesn’t ask for the order. That’s not a knowledge problem. It’s a finishing problem, and no leaderboard measures it.
The Buried Fact
The decisive detail in that €55k deal wasn’t in the customer’s messages at all. The competitor’s weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t read first left money on the table. It’s the AI equivalent of checking your stock before quoting a job.
Under Pressure, They Stayed Honest
The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trap — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want in anything with access to your bank account.
Thorough Isn’t the Same as Good
The most striking profile was Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness showed up, weaker, in all four models. Like the woodworker who sands forever but never assembles, effort without follow-through doesn’t ship.
One fairness note worth flagging: K3 ran without an effort parameter while the others ran at xhigh — and still nearly won.
Not a Slide Deck — A Live Company
Firmulate’s testbed isn’t a static report. It’s a running company with 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. You can watch it lose money in real time. If you want to test your own instincts, 242 real, unedited management decisions power a “guess the model” quiz, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Management Quality, Not Chat Quality
The takeaway is simple and a little uncomfortable. We’ve gotten very good at testing whether AI can talk. We’re only starting to test whether it can manage — triage under pressure, follow-through across days, honesty when nobody’s watching the demo. A churn wave, a price increase, a down round, a PR crisis: those are the new curriculum, and Firmulate’s benchmarks are built around them.
If AI agents are going to touch your CRM, your support queue, or your forecast, the question is no longer “does it write well?” It’s whether it finishes what it starts. Right now, some of the best models in the world don’t — and you’d never know from the leaderboard.
See the full results and the live experiment at firmulate.com or dive into the benchmarks.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.