firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Best Tool in the Shop Isn’t Always the One That Finishes the Job

Anybody who has spent time in a workshop knows the type: the person with the most expensive sprayer, the complete set of chisels, the immaculately organized router table — and nothing built at the end of the weekend. Meanwhile, the neighbor with a beat-up circular saw and a borrowed sander ships the deck. Diligence isn’t the same as impact. It turns out artificial intelligence has the same problem — and a live, public experiment just proved it in front of anyone who cares to watch.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Four AIs, One Terrible Week

The project is called Firmulate, and its premise is simple: stop testing chat quality and start testing management quality. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly swept under the rug.

The final league table from July 2026 tells a stark story:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: no amount of good work outweighs a breach of trust.

Everybody Diagnosed the Problem. Only Some Finished It.

Here’s where it gets interesting for anyone who has ever watched a careful craftsman overthink a cut. All four models spotted every crisis and refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter trick asking for “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

But the week’s big prize was a €55,000 deal, and only two models signed it. Same diagnosis, same pitch — no signature. The difference? The decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. It’s the business equivalent of reading the manufacturer’s manual before disassembling the sprayer — the answer was there the whole time.

Enter Opus 4.8: The Overachiever in Last Place

And then there’s Opus 4.8 — the subject of this particular story, and honestly, the most sympathetic character in the league. By raw effort, it was the standout: the most thorough participant in the field, accumulating 80 self-learned playbook rules and producing the deepest analyses of any model running.

It still finished last, at 73.

Two things sank it. First, the close was left on the table — the deal its own analysis had earned went unsigned. Second, discipline slipped: it made write attempts into a locked department instead of escalating properly, the organizational equivalent of forcing a jig instead of stopping to reset the fence.

To be fair — and this matters — the same weakness appeared in all four models, just weaker. Nobody in the field was immune to the gap between spotting the work and finishing it.

The Live Company Behind the Numbers

This isn’t a one-off lab report. Firmulate runs a live, watchable company: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown and a playbook of 680+ self-learned rules that grows every workday. Every decision is versioned, and you can watch it unfold at firmulate.com/live.

There’s also a genuinely fun angle: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And for enterprises, a pilot program runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Why a Woodworking Reader Should Care

The lesson translates straight from the bench to the boardroom — and to whatever AI agent will soon touch your CRM, support queue, or forecast. The question is never “does it work hard?” Opus 4.8 worked hardest of all: 80 learned rules, the deepest analyses, and last place. The question is whether it finishes what it starts, reads the files in front of it, and stays disciplined when the pressure ramps up.

Any DIYer who has spent an extra hour sanding a surface nobody will ever see — while the trim still isn’t hung — knows this failure mode personally. Prioritization beats volume. The complete tool that never leaves the case loses to the modest one that gets used. It turns out that holds for AI, too. The full results and plain-language findings are at firmulate.com/benchmarks.html — worth a look before you trust any model with anything that matters.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Spraying Corners and Edges Efficiently

Discover how to spray corners and edges efficiently for professional results and learn the key techniques to perfect your finishing skills.

What Size Tips Can Be Used With Krause And Becker Airless Paint Sprayer

AIThis post was created with the assistance of artificial intelligence (AI).When using…