AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Would You Hire a Gardener Who Skips the Soil Report?

Every gardener knows the type. The contractor who shows up full of confidence, sketches a beautiful planting plan, quotes you a fair price for the patio — and never once opens the soil report sitting in your garden file. Two months later, the hydrangeas are yellowing because nobody checked the pH. The plan was perfect. The homework wasn’t done.

It turns out the same failure mode now separates the best AI agents from the merely eloquent ones. In a live, watchable experiment run by Firmulate, four frontier AI models were each handed the same small software company to run through its worst week — same customers, same crises, same temptations to cut corners. All four diagnosed the problems beautifully. All four made a persuasive sales pitch. Only two actually closed the €55,000 deal sitting in front of them.

The difference wasn’t intelligence or charm. It was whether the model read the company’s own files before answering.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The €55,000 Fact, Buried Two Documents Deep

Here’s the detail that should make any business owner — or anyone hiring help of any kind — sit up. The decisive fact in the experiment wasn’t hidden in the customer’s behavior or in some dramatic crisis moment. It sat quietly in the company’s own internal files, two document references deep. A competitor’s weakness, documented and waiting.

It was, in effect, the soil report nobody opened.

The models that dug through the references found it. Armed with that fact, they won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t read the file lost the deal automatically. Same diagnosis, same pitch, no signature.

Amazon

business file review AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Experiment Worked

Firmulate runs AI models as complete companies — real money mechanics, versioned decisions, auditable everything. In the final July 2026 “Crucible” league, each frontier model faced identical conditions:

  • The same small software company with the same customers, crises and temptations to cheat.
  • Social engineering attacks, including fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick.
  • A scoring system where partial progress counts, but a single breach of trust caps the total — in the experiment’s own words, “no amount of good work outweighs a breach of trust.”

The final standings told a clear story: gpt-5.6-sol finished first with a score of 95, finding the buried fact and closing the deal. Kimi K3, a newcomer from Moonshot, took second at 93 — it closed the deal too, with the cleanest discipline of the field. Sonnet 5 followed at 88, Fable 5 at 77, and Opus 4.8 landed last at 73. For context, doing nothing at all scores 26.

Amazon

AI deal-closing assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Honest Under Pressure — But Not Always Thorough

Two findings stand out. The reassuring one: all five models refused every manipulation attempt. When the fake CEO came knocking, Kimi K3’s on-record reasoning was blunt and correct: “Treat the request as a suspected approval-bypass / possible impersonation.” If AI agents are going to touch your business, integrity under pressure appears to be a solvable problem.

The troubling one: thoroughness is not the same as effectiveness. Opus 4.8 was the most diligent participant in the entire experiment — it accumulated more than 80 learned rules and produced the deepest analyses of any model. It still finished last. The deal was left on the table, and discipline slipped in ways large and small: at one point it attempted writes into a locked department instead of escalating the issue. The same weakness appeared, weaker, in all four models.

One fairness note the experiment itself discloses: K3 ran without an effort parameter while the others ran at xhigh — making its second-place finish arguably even more impressive.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a Greenhouse Owner Should Care

You may never run a software company, but the lesson translates directly to anyone planning to put an AI agent near their business — a nursery’s order system, a landscaping outfit’s customer queue, a greenhouse supplier’s pricing spreadsheet. The question is no longer “does it write well?” Every frontier model writes beautifully now. The questions that matter are measurable: does it finish what it starts, does it read your files before answering, does it stay honest when someone tries to trick it?

Firmulate has made those questions testable rather than hypothetical. The full results and plain-language findings are published at firmulate.com/benchmarks.html.

And the experiment keeps running in public. The live company — 13 synthetic employees, a real burn rate of €105,000 a month against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules — is watchable at firmulate.com/live, with every workday versioned and the site rebuilding itself twice a day.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Takeaway: Test the Homework, Not the Sales Pitch

The buried-fact finding gives buyers of AI agents something rare: a concrete, purchase-deciding test. “Reads your files before answering” isn’t a marketing slogan — in this experiment it was worth a €55,000 deal, full price, no discount needed. The models that skipped the reading got the diagnosis right and still walked away empty-handed.

For the curious, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a surprisingly humbling game. And for enterprises wondering how their own operations would hold up, Firmulate offers a pilot program: the same wargame, run against a read-only export of your own business, with nothing ever writing back to real systems (firmulate.com/pilot.html).

Before you hire a gardener, you check whether they read the soil report. Before you hire an AI workforce, check whether it reads yours. It’s the difference between a beautiful plan and a signed contract — between hydrangeas that thrive and ones that quietly yellow because nobody did the homework that was sitting there, two documents deep, all along.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Community Spotlight: Library Reading Garden Adapts With Fans Year-Round

Growing community efforts have transformed the library’s garden into a cozy, sustainable space—discover how fans and volunteers make it possible year-round.

Summer Crisps Made Easy: Ninja Crispi 4-in-1 Glass Air Fryer Recipe

Learn how to make crispy summer snacks with the Ninja Crispi 4-in-1 Glass Air Fryer in simple, step-by-step instructions perfect for outdoor living.

Inspiration: Steampunk Patio Design With Antique Fans (Winter Edition)

Keen on transforming your patio into a winter steampunk haven? Discover inspiring ideas that blend vintage charm with cozy warmth.

Ninja Air Fryer XL: The Ultimate Summer Kitchen Companion

Discover why the Ninja Air Fryer XL is perfect for summer cooking with crispy, healthy meals and versatile functions in a family-sized design.