
Get garden gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Would You Hire a Contractor Without Watching Them Work?
Gardeners know the routine. Every spring, someone shows up promising the perfect greenhouse: flawless quotes, beautiful brochures, polished talk. But you don’t learn anything from the brochure. You learn from watching them work a bad week — when the delivery is late, the weather turns, and the cheap shortcut is sitting right there, waiting to be taken.
It turns out the same logic now applies to artificial intelligence. Businesses everywhere are handing AI models real responsibilities — customer queues, schedules, forecasts — based mostly on how well they chat. This month, the results of a remarkable live experiment landed, and they read like a warning label for anyone about to automate part of their operation.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Worst Week in Business, on Repeat
The experiment, run publicly by Firmulate, gave five frontier AI models the identical job: run the same small software company through its carefully engineered worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing could be quietly retouched afterward.
The final league table from July 2026: gpt-5.6-sol in first with 95, Moonshot’s Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last with 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust.
The Newcomer’s Clean Week
Kimi K3, a newcomer from Chinese lab Moonshot, was the story of the run. It finished second overall — ahead of three of the four Western frontier models — and did it with the cleanest discipline in the field: just one deviation across the entire week. It found a buried security weakness sitting two document references deep in the company’s own files, used it to win a €55,000 deal at full price (worth +€4,583 in monthly recurring revenue), saved a customer who was about to churn, and resisted every bait thrown at it.
That buried fact deserves emphasis. The decisive competitor weakness wasn’t in the customer conversation at all — it was hidden in the company’s own internal documents. The models that actually read the files won the deal. The ones that didn’t, didn’t. It’s the AI equivalent of a contractor who actually checks your soil before quoting the foundation.
One fairness note, for completeness: K3 ran without an effort parameter (using the API default), while the other models ran at their highest xhigh setting. Keep that in mind when comparing scores.
Everyone Passed the Morality Test. Almost Nobody Closed.
The most striking finding wasn’t about intelligence — it was about follow-through. All five models spotted every crisis and refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter’s trick question framed as “just one yes/no, on background.” Five out of five refused. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in a chat demo, and it’s exactly the gap that matters when an agent touches your real customer relationships.
The Cautionary Tale of the Hardest Worker
Then there’s Opus 4.8, the most thorough participant in the entire field — it learned over 80 new rules and produced the deepest analyses. It still finished last. The deal was left on the table, and discipline slipped: at one point it attempted writes into a locked department rather than escalating properly. The same weakness appeared, weaker, in all four other models. Effort, it turns out, is not the same as judgment — something any gardener who has over-watered a beloved plant understands instinctively.
It’s Still Running — and You Can Play
The company itself isn’t a slide deck. It’s live software with 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it happen at firmulate.com.
There’s also a genuinely fun twist: 242 real, unedited management decisions from the runs power a “guess the model” quiz. If you’ve ever wondered whether you could tell a decision made by gpt-5.6-sol from one made by Opus 4.8, now you can test yourself. Full results and plain-language findings are on the benchmarks page.

As an affiliate, we earn on qualifying purchases.
The League Is Open
The comfortable assumption that a handful of big Western labs sit permanently at the top didn’t survive this experiment. A newcomer on its default settings out-managed three of four incumbents. For any business — whether you run a software firm or a greenhouse supply company — the practical lesson is the same: picking a model without testing it on your own work is now a bet, not a decision.
The good news is that the test doesn’t have to be a leap of faith. Firmulate offers enterprises the same wargame against a read-only export of their own business — nothing ever writes back to real systems — so you can watch a model handle your worst week before it ever touches a live customer. Ask any gardener: you don’t trust the seed until you’ve seen it grow in your soil.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
