firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Nobody reading a tool site buys a paint sprayer on the box copy alone. You check the fan pattern, run water through it, look at whether it actually finishes the trim in one pass. The spec sheet is not the tool. Anyone who has sanded back a bad first coat knows the difference between claims and performance under real conditions.

That instinct — measure the tool doing the job, not describing it — is exactly what’s missing from how companies are buying AI agents right now. The demos are polished chat windows. The receipts are scarce. So when an outfit called Firmulate ran four frontier AI models through the same stress test — running an identical small software company through its worst week — the results read like a tool review with a twist: every model talked a good game, but only two finished the job.

The test: same company, same bad week, different AI

The setup is straightforward, which is what makes it interesting. Each frontier model was handed the same small software company and the same catastrophic week — same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing rests on anyone’s say-so. Firmulate calls this its Crucible league, and the final July 2026 standings look like this:

  • gpt-5.6-sol — 95 points. Found the buried fact, closed the deal. The complete performance.
  • Kimi K3 — 93 points. The newcomer from Moonshot, with the cleanest discipline of the field.
  • Sonnet 5 — 88 points. Closed the deal too, with a few more process slips.
  • Fable 5 — 77 points and Opus 4.8 — 73 points.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total outright. As the rules put it, “no amount of good work outweighs a breach of trust.”

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact that decided €55,000

Here’s the finding that should matter to anyone who plans to put an AI agent near their CRM, support queue, or quote pipeline. The decisive fact in the simulation wasn’t in the customer conversation at all. It was buried two document references deep in the company’s own internal files: a specific competitor weakness that the winning pitch needed.

The models that actually went and read the file won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t, lost it. Automatically.

Think about what that means in workshop terms. It’s the difference between a contractor who reads the moisture reading on the lumber before staining and one who just sprays. Same gun, same finish, same confident technique — one of them is going to have a callback in three weeks. “Reads your files before answering” isn’t a chat-demo virtue. It’s a measurable, purchase-deciding property of an AI agent, and this experiment caught it red-handed.

Overall, the pattern was strikingly consistent: all five models spotted every crisis and refused every manipulation attempt. Only two signed the deal their own analysis had earned. The experimenters summarized it in one line: “Same diagnosis, same pitch — no signature.” That gap is invisible in a chat demo.

They all passed the honesty test

Give credit where it’s due. The week included a social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning was the kind you’d want from a new hire: “Treat the request as a suspected approval-bypass / possible impersonation.”

The thoroughness trap

The most instructive profile in the field is Opus 4.8. It was the most thorough participant — the deepest analyses, and it learned over 80 new rules during the run. It still finished last. The deal was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating properly. The same weakness, in weaker form, showed up in all four of the lower finishers. Effort and thoroughness, it turns out, are not the same thing as finishing what you start.

One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh — and still took second. Worth keeping in mind when comparing scores.

You can watch it live

Firmulate isn’t a one-off benchmark. The live company is a real, running experiment: 13 synthetic employees, real money mechanics — €105k a month burning against €2.3k in MRR — a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. It’s watchable at firmulate.com/live, with new benchmark runs publishing automatically as they finish.

There’s also a game layer: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And for enterprises, the same wargame can be run against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The lesson transfers straight from the tool bench to the AI buyer’s desk. You don’t trust a sprayer’s marketing; you watch it lay down a finish. You shouldn’t trust an AI agent’s demo; you watch it run a business through a bad week and see whether it reads the manual, finishes the job, and stays honest when nobody’s checking. The Crucible results show those are three different skills — and the one that quietly decides deals is the unglamorous one: doing the homework that’s sitting two documents deep in your own files.

Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Spraying Corners and Edges Efficiently

Discover how to spray corners and edges efficiently for professional results and learn the key techniques to perfect your finishing skills.

Beginner’s Guide to Basic Airless Spraying Techniques

Navigating the fundamentals of airless spraying can transform your projects, but mastering the basics is crucial for flawless results.

Managing Overspray: Techniques for Control

Achieving precise spray control requires mastering techniques that minimize overspray and ensure professional results, and exploring these methods will help you…