AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garden gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A Lesson From the Greenhouse: Partial Credit Is Real Credit

Any gardener knows you don’t get a zero just because the tomatoes came in late. You amended the soil, you watered, you kept the blight at bay — the season wasn’t a total loss. Judging growing things, like judging people (or AI), on a pass/fail basis tells you almost nothing. So when an AI benchmark called the Crucible League recently published its final standings, one design choice stood out to anyone who has ever coaxed a stubborn crop through a bad summer: a manager that does nothing at all still scores 26 points out of 100. Not zero. Twenty-six.

That floor isn’t a bug or grade inflation. It’s the whole philosophy of the test — and it says a lot about what honest evaluation of AI agents should look like.

Amazon

AI decision-making benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Worst Week a Small Company Ever Had

Firmulate, a public project that runs AI models as complete companies, gave four frontier models the same impossible job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing could be quietly rewritten after the fact.

The final July 2026 league table reads: gpt-5.6-sol in first with 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 in last at 73.

Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Doing Nothing Scores 26

Here’s the reasoning, and it maps neatly onto gardening logic. Even a manager who takes no decisive action still does some things right in a company that’s already functioning: crises get noticed, obvious mistakes get avoided, the business doesn’t burn down. Partial progress counts, because in the real world partial progress is still value — the half-watered garden still out-produves the abandoned one.

The flip side is sterner. A single breach of trust caps the total grade entirely. As the benchmark’s own framing puts it: “no amount of good work outweighs a breach of trust.” A brilliant manager who lies once is not a brilliant manager with a small deduction. That’s a rule most gardeners would recognize: one contaminated compost batch can ruin the whole bed.

Amazon

AI performance testing platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

And Why Nobody Scores a Round 100

Notice the top score is 95, not 100. The benchmark’s designers treat a perfect round number as a warning sign, not a triumph — the same instinct that makes an experienced grower suspicious of a tomato plant with zero blemishes. Something got missed, or the test was too easy. Distrust of tidy 100s is baked into the culture.

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Separated the Winners

The striking finding: all five models spotted every crisis and refused every manipulation attempt — including fake CEO messages that escalated over three stages, plus a reporter’s “just one yes/no, on background” trick. Five out of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The difference? The decisive competitor weakness was buried two document references deep in the company’s own files, not in the customer event. The models that actually read their own file cabinet won the deal at full price, worth +€4,583 in monthly recurring revenue. In gardening terms: the ones that tested their own soil before planting.

The Thoroughness Paradox

Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Effort spent isn’t the same as work finished.

One Fairness Footnote

K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won. That transparency is part of the same honest-benchmark ethos.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

You Can Watch It Grow

The experiment isn’t a one-off paper. Firmulate runs a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. There’s also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business.

The lesson for anyone who manages anything — a company, a greenhouse, a garden: reward partial progress, never forgive a breach of trust, read your own files before you act, and be suspicious of anything that scores a perfect 100.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Host the Ultimate Summer Pool Party with Ninja Foodi Pro 5-in-1 Grill

Make your summer pool party effortless with the Ninja Foodi Pro 5-in-1 Indoor Grill — versatile, smoke-free, and perfect for grilling, air frying, and more.

Summer Recipe Ideas with the Ninja Foodi 10 Qt Air Fryer

Discover easy, fresh summer recipes using the Ninja Foodi 10 Quart DualZone XL Air Fryer for quick, versatile outdoor cooking.

How to Document Your Patio Makeover for Social Media

Aiming to showcase your patio transformation? Discover the key steps to capturing and sharing your makeover effectively.

Madrid Launches Guide To Climate-Friendly Schoolyards

Madrid promotes climate transformation of schoolyards with a new action guide, aiming to improve urban resilience and environmental education.