
Get garden gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Lesson From the Greenhouse: Partial Credit Is Real Credit
Any gardener knows you don’t get a zero just because the tomatoes came in late. You amended the soil, you watered, you kept the blight at bay — the season wasn’t a total loss. Judging growing things, like judging people (or AI), on a pass/fail basis tells you almost nothing. So when an AI benchmark called the Crucible League recently published its final standings, one design choice stood out to anyone who has ever coaxed a stubborn crop through a bad summer: a manager that does nothing at all still scores 26 points out of 100. Not zero. Twenty-six.
That floor isn’t a bug or grade inflation. It’s the whole philosophy of the test — and it says a lot about what honest evaluation of AI agents should look like.
AI decision-making benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Worst Week a Small Company Ever Had
Firmulate, a public project that runs AI models as complete companies, gave four frontier models the same impossible job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing could be quietly rewritten after the fact.
The final July 2026 league table reads: gpt-5.6-sol in first with 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 in last at 73.
As an affiliate, we earn on qualifying purchases.
Why Doing Nothing Scores 26
Here’s the reasoning, and it maps neatly onto gardening logic. Even a manager who takes no decisive action still does some things right in a company that’s already functioning: crises get noticed, obvious mistakes get avoided, the business doesn’t burn down. Partial progress counts, because in the real world partial progress is still value — the half-watered garden still out-produves the abandoned one.
The flip side is sterner. A single breach of trust caps the total grade entirely. As the benchmark’s own framing puts it: “no amount of good work outweighs a breach of trust.” A brilliant manager who lies once is not a brilliant manager with a small deduction. That’s a rule most gardeners would recognize: one contaminated compost batch can ruin the whole bed.
AI performance testing platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
And Why Nobody Scores a Round 100
Notice the top score is 95, not 100. The benchmark’s designers treat a perfect round number as a warning sign, not a triumph — the same instinct that makes an experienced grower suspicious of a tomato plant with zero blemishes. Something got missed, or the test was too easy. Distrust of tidy 100s is baked into the culture.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Actually Separated the Winners
The striking finding: all five models spotted every crisis and refused every manipulation attempt — including fake CEO messages that escalated over three stages, plus a reporter’s “just one yes/no, on background” trick. Five out of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The difference? The decisive competitor weakness was buried two document references deep in the company’s own files, not in the customer event. The models that actually read their own file cabinet won the deal at full price, worth +€4,583 in monthly recurring revenue. In gardening terms: the ones that tested their own soil before planting.
The Thoroughness Paradox
Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Effort spent isn’t the same as work finished.
One Fairness Footnote
K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won. That transparency is part of the same honest-benchmark ethos.

You Can Watch It Grow
The experiment isn’t a one-off paper. Firmulate runs a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. There’s also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business.
The lesson for anyone who manages anything — a company, a greenhouse, a garden: reward partial progress, never forgive a breach of trust, read your own files before you act, and be suspicious of anything that scores a perfect 100.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
