
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Best Tool in the Shop Isn’t Always the One That Finishes the Job
Anybody who has spent time in a workshop knows the type: the person with the most expensive sprayer, the complete set of chisels, the immaculately organized router table — and nothing built at the end of the weekend. Meanwhile, the neighbor with a beat-up circular saw and a borrowed sander ships the deck. Diligence isn’t the same as impact. It turns out artificial intelligence has the same problem — and a live, public experiment just proved it in front of anyone who cares to watch.
As an affiliate, we earn on qualifying purchases.
The Experiment: Four AIs, One Terrible Week
The project is called Firmulate, and its premise is simple: stop testing chat quality and start testing management quality. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly swept under the rug.
The final league table from July 2026 tells a stark story:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: no amount of good work outweighs a breach of trust.
Everybody Diagnosed the Problem. Only Some Finished It.
Here’s where it gets interesting for anyone who has ever watched a careful craftsman overthink a cut. All four models spotted every crisis and refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter trick asking for “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
But the week’s big prize was a €55,000 deal, and only two models signed it. Same diagnosis, same pitch — no signature. The difference? The decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. It’s the business equivalent of reading the manufacturer’s manual before disassembling the sprayer — the answer was there the whole time.
Enter Opus 4.8: The Overachiever in Last Place
And then there’s Opus 4.8 — the subject of this particular story, and honestly, the most sympathetic character in the league. By raw effort, it was the standout: the most thorough participant in the field, accumulating 80 self-learned playbook rules and producing the deepest analyses of any model running.
It still finished last, at 73.
Two things sank it. First, the close was left on the table — the deal its own analysis had earned went unsigned. Second, discipline slipped: it made write attempts into a locked department instead of escalating properly, the organizational equivalent of forcing a jig instead of stopping to reset the fence.
To be fair — and this matters — the same weakness appeared in all four models, just weaker. Nobody in the field was immune to the gap between spotting the work and finishing it.
The Live Company Behind the Numbers
This isn’t a one-off lab report. Firmulate runs a live, watchable company: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown and a playbook of 680+ self-learned rules that grows every workday. Every decision is versioned, and you can watch it unfold at firmulate.com/live.
There’s also a genuinely fun angle: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And for enterprises, a pilot program runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Why a Woodworking Reader Should Care
The lesson translates straight from the bench to the boardroom — and to whatever AI agent will soon touch your CRM, support queue, or forecast. The question is never “does it work hard?” Opus 4.8 worked hardest of all: 80 learned rules, the deepest analyses, and last place. The question is whether it finishes what it starts, reads the files in front of it, and stays disciplined when the pressure ramps up.
Any DIYer who has spent an extra hour sanding a surface nobody will ever see — while the trim still isn’t hung — knows this failure mode personally. Prioritization beats volume. The complete tool that never leaves the case loses to the modest one that gets used. It turns out that holds for AI, too. The full results and plain-language findings are at firmulate.com/benchmarks.html — worth a look before you trust any model with anything that matters.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.