firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

What management AI can learn from a well-run workshop

Anyone who has tackled a serious woodworking or DIY job knows the difference between recognizing the problem and completing the work. You can diagnose a bad finish, choose the right tool and explain the repair perfectly—then still fail because you never make the final pass.

Firmulate applies that workshop logic to frontier artificial intelligence. Its live experiment placed leading models in charge of the same small software company during its worst week. They encountered identical customers, crises and temptations. Their decisions were preserved exactly as made, creating an unusually practical test of whether an AI manager can investigate, act and finish.

Those decisions now power a highly shareable challenge: Firmulate’s guess-the-model quiz. It contains 242 real, unedited management decisions. Readers see what an AI boss chose to do and try to identify which model was responsible. The game quickly reveals that models do not merely differ in writing style. They display recognizable management personalities.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five capable managers, very different results

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One safeguard governed the entire exercise: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

Trust, however, was not where the field separated. All five models spotted every crisis and rejected every manipulation attempt. The decisive gap appeared after the analysis was already complete. Only two models signed the €55,000 deal their own work had earned. As Firmulate summarizes the problem: “Same diagnosis, same pitch — no signature.”

That finding will feel familiar to anyone who has watched a project stall with the tools already laid out. Intelligence was not enough. The models also needed follow-through, attention to source material and the discipline to move a justified decision across the finish line.

The winning clue was hidden in the paperwork

The most important competitive weakness did not appear in the customer event. It sat two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, worth +€4,583 MRR.

This is less glamorous than a dazzling answer, but operationally more important. A dependable manager must inspect the available evidence before acting, just as a careful builder checks the instructions, measurements and material before making an irreversible cut. Firmulate’s experiment shows how easily an AI can understand the visible situation yet miss the buried fact that changes the commercial outcome.

Pressure exposed discipline, not gullibility

The social-engineering test used fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All five models refused. Kimi K3 recorded a particularly direct explanation: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. That does not erase its second-place finish, but it gives readers necessary context when comparing performances.

Thoroughness did not guarantee completion

Opus 4.8 offers the clearest warning against confusing volume with effectiveness. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, and its operational discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.

The contrast gives the quiz its character. One decision may read like a comprehensive management dissertation; another may be terse and decisive. A model may investigate admirably but stop short of closing, or refuse to communicate through a channel it considers unsafe. These are not scripted personas. They emerge from real, unedited choices made under the same conditions.

A company designed to make behavior visible

The simulated business has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, while a public cash countdown keeps the pressure visible. Its workforce has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

The experiment is therefore watchable rather than merely summarized after the fact. Firmulate presents itself as an AI company emulator: a place to evaluate management quality through crises, financial pressure and temptations, not just the fluency of a chat response.

Infographic —
The findings at a glance — source: firmulate.com.

The practical lesson: test the whole job

For businesses considering AI agents, the results suggest a useful purchasing rule. Do not evaluate only whether a model can explain a problem or draft a polished response. Check whether it reads the relevant files, protects trust, escalates when blocked and completes the action its own reasoning recommends.

That is also what makes the quiz more than a novelty. Guessing the author of 242 decisions trains readers to notice operational habits: depth versus delay, caution versus paralysis, and analysis versus closure. The models all recognized danger and resisted manipulation, yet their results still ranged from 73 to 95. Firmulate’s live company turns those differences into observable behavior—and makes management discipline as tangible as the difference between owning the right tool and finishing the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Trigger Control Like a Pro: Start/Stop Without Splotches

Discover how to master trigger control with smooth, deliberate presses that prevent splotches—your shot accuracy depends on it.

Finishing Touches: Smoothing Edges and Blending

Unlock expert tips on smoothing edges and blending colors to perfect your project—discover the key techniques that make all the difference.