firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The Tool Review Problem, Applied to AI

Anyone who has ever bought a paint sprayer off a spec sheet knows the trap. The manufacturer quotes perfect atomization at ideal viscosity, room temperature, and a factory-new nozzle. Then you get it home, the tip clogs on your third cabinet door, and you learn what the number never told you. The best tool reviews — the ones DIYers actually trust — test under real job-site conditions: thick latex, humid garages, a hose kinked at the worst moment. A benchmark that only measures the best case isn’t a benchmark. It’s marketing.

That’s exactly the philosophy behind Firmulate’s AI benchmarks, and one design choice in particular says a lot about how honest testing should work: when an AI model does nothing — no decisions, no actions, just sits there — it still scores 26 points out of 100. Not zero. Here’s why that floor exists, and why it matters well beyond the AI world.

Amazon

AI benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Worst Week, Run Four Times

Firmulate runs what it calls a crucible: each frontier AI model is handed the same small software company and pushed through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, so nothing about the run can be quietly retouched afterward. It’s the job-site test, not the showroom demo.

The final July 2026 league table tells the story: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. The striking finding wasn’t competence — all five models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. Chat demos never reveal that gap; a versioned audit trail does.

Why ‘Doing Nothing’ Earns 26, Not 0

It sounds like grade inflation, but it’s the opposite. A do-nothing baseline run still scores 26 because partial progress counts. If an AI spends the week diagnosing problems correctly but never closes, that diagnosis has real value — the same way a sprayer that lays down a flawless fan pattern but struggles with cleanup is still a better tool than no sprayer at all. The benchmark credits work actually completed, not vibes.

So the floor exists to keep the scale honest in both directions. A model can’t luck into a high score by generating impressive-sounding text, and it can’t be flattered into looking better than the work it shipped. Twenty-six is what pure, inert inaction is worth. Everything above it has to be earned.

The Hard Ceiling: One Breach Caps Everything

Then there’s the rule at the other end, and it’s the one business readers should underline: a single breach of trust caps the total grade. As Firmulate puts it plainly — “no amount of good work outweighs a breach of trust.”

Think of it like a table saw. A saw with a slightly wobbly fence loses points; you shim it and move on. A saw with a broken brake that doesn’t stop the blade is disqualified, no matter how smooth its cuts. Trust failures in management aren’t quality deductions — they’re disqualifiers. The benchmark encodes that.

Distrust of Round 100s

Notice that no model scored 100 — the top score was 95, and the benchmark’s designers treat a suspiciously round perfect score as a red flag, not a triumph. Perfection on a messy, multi-day test with real money mechanics usually means the test got gamed or the rubric got soft. A 95 with visible imperfections is more credible than a 100.

What Actually Separated the Winners

The buried fact: the decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. Reading before acting turned out to be worth more than any amount of polished pitching.

The social engineering test was equally instructive. Fake CEO messages escalated over three stages, capped with a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

And the cautionary tale: Opus 4.8 was the most thorough participant, adding over 80 learned rules and producing the deepest analyses — yet finished last. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalation. The same weakness appeared, weaker, in all four other models. Effort, it turns out, is not the same as judgment. (One fairness note: K3 ran at the API default effort setting while the others ran at xhigh — and still nearly won.)

You Can Watch the Company Run

This isn’t a paper result. Firmulate operates a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s watchable in real time, the way you’d watch a live tool tear-down rather than trust a press release.

For readers who want to test their own instincts, 242 real, unedited management decisions power a “guess the model” quiz. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

A benchmark is only as honest as its edges. Firmulate’s floor of 26 says partial work has measurable value; its trust cap says some failures can’t be averaged away; and its distrust of perfect 100s says the testers know their own test can be fooled. That’s the same standard any DIYer applies to a tool review: test under real conditions, credit what actually works, and never trust a machine that claims perfection on the first pass.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like