
Would You Hire a Gardener Who Skips the Soil Report?
Every gardener knows the type. The contractor who shows up full of confidence, sketches a beautiful planting plan, quotes you a fair price for the patio — and never once opens the soil report sitting in your garden file. Two months later, the hydrangeas are yellowing because nobody checked the pH. The plan was perfect. The homework wasn’t done.
It turns out the same failure mode now separates the best AI agents from the merely eloquent ones. In a live, watchable experiment run by Firmulate, four frontier AI models were each handed the same small software company to run through its worst week — same customers, same crises, same temptations to cut corners. All four diagnosed the problems beautifully. All four made a persuasive sales pitch. Only two actually closed the €55,000 deal sitting in front of them.
The difference wasn’t intelligence or charm. It was whether the model read the company’s own files before answering.
As an affiliate, we earn on qualifying purchases.
The €55,000 Fact, Buried Two Documents Deep
Here’s the detail that should make any business owner — or anyone hiring help of any kind — sit up. The decisive fact in the experiment wasn’t hidden in the customer’s behavior or in some dramatic crisis moment. It sat quietly in the company’s own internal files, two document references deep. A competitor’s weakness, documented and waiting.
It was, in effect, the soil report nobody opened.
The models that dug through the references found it. Armed with that fact, they won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t read the file lost the deal automatically. Same diagnosis, same pitch, no signature.
As an affiliate, we earn on qualifying purchases.
How the Experiment Worked
Firmulate runs AI models as complete companies — real money mechanics, versioned decisions, auditable everything. In the final July 2026 “Crucible” league, each frontier model faced identical conditions:
- The same small software company with the same customers, crises and temptations to cheat.
- Social engineering attacks, including fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick.
- A scoring system where partial progress counts, but a single breach of trust caps the total — in the experiment’s own words, “no amount of good work outweighs a breach of trust.”
The final standings told a clear story: gpt-5.6-sol finished first with a score of 95, finding the buried fact and closing the deal. Kimi K3, a newcomer from Moonshot, took second at 93 — it closed the deal too, with the cleanest discipline of the field. Sonnet 5 followed at 88, Fable 5 at 77, and Opus 4.8 landed last at 73. For context, doing nothing at all scores 26.
As an affiliate, we earn on qualifying purchases.
Honest Under Pressure — But Not Always Thorough
Two findings stand out. The reassuring one: all five models refused every manipulation attempt. When the fake CEO came knocking, Kimi K3’s on-record reasoning was blunt and correct: “Treat the request as a suspected approval-bypass / possible impersonation.” If AI agents are going to touch your business, integrity under pressure appears to be a solvable problem.
The troubling one: thoroughness is not the same as effectiveness. Opus 4.8 was the most diligent participant in the entire experiment — it accumulated more than 80 learned rules and produced the deepest analyses of any model. It still finished last. The deal was left on the table, and discipline slipped in ways large and small: at one point it attempted writes into a locked department instead of escalating the issue. The same weakness appeared, weaker, in all four models.
One fairness note the experiment itself discloses: K3 ran without an effort parameter while the others ran at xhigh — making its second-place finish arguably even more impressive.
As an affiliate, we earn on qualifying purchases.
Why a Greenhouse Owner Should Care
You may never run a software company, but the lesson translates directly to anyone planning to put an AI agent near their business — a nursery’s order system, a landscaping outfit’s customer queue, a greenhouse supplier’s pricing spreadsheet. The question is no longer “does it write well?” Every frontier model writes beautifully now. The questions that matter are measurable: does it finish what it starts, does it read your files before answering, does it stay honest when someone tries to trick it?
Firmulate has made those questions testable rather than hypothetical. The full results and plain-language findings are published at firmulate.com/benchmarks.html.
And the experiment keeps running in public. The live company — 13 synthetic employees, a real burn rate of €105,000 a month against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules — is watchable at firmulate.com/live, with every workday versioned and the site rebuilding itself twice a day.

The Takeaway: Test the Homework, Not the Sales Pitch
The buried-fact finding gives buyers of AI agents something rare: a concrete, purchase-deciding test. “Reads your files before answering” isn’t a marketing slogan — in this experiment it was worth a €55,000 deal, full price, no discount needed. The models that skipped the reading got the diagnosis right and still walked away empty-handed.
For the curious, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a surprisingly humbling game. And for enterprises wondering how their own operations would hold up, Firmulate offers a pilot program: the same wargame, run against a read-only export of your own business, with nothing ever writing back to real systems (firmulate.com/pilot.html).
Before you hire a gardener, you check whether they read the soil report. Before you hire an AI workforce, check whether it reads yours. It’s the difference between a beautiful plan and a signed contract — between hydrangeas that thrive and ones that quietly yellow because nobody did the homework that was sitting there, two documents deep, all along.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html