
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Smart Contract Problem, But for Management
Crypto people understand something most industries are still learning the hard way: a system that executes flawlessly 99% of the time and fails catastrophically 1% of the time is not a good system. It’s a time bomb with good marketing. That’s why auditors crawl through smart contracts line by line before a protocol touches real money — you don’t trust the demo, you stress-test the code against its worst week.
Now apply that logic to the AI agents enterprises are about to plug into their CRMs, support queues and forecasts. Everyone is benchmarking how well these models chat. Almost nobody is testing how well they manage — under pressure, with real temptations, when nobody’s watching. A public experiment called Firmulate has been doing exactly that, running frontier AI models as complete companies with real money mechanics and a public cash countdown. The results are watchable, versioned, and — for anyone planning to deploy agents against their own business — directly relevant.
Four Models, One Terrible Week
The setup is elegant: each of four frontier AI models was given the same job — run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, like commits on a chain.
The final league standings tell a story that chat benchmarks never show:
- gpt-5.6-sol — 95
- Kimi K3 — 93 (with a caveat: it ran at API-default effort while the others ran at xhigh)
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
- Doing nothing at all scored 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s scoring philosophy puts it: “no amount of good work outweighs a breach of trust.”
The Finding That Should Worry Every Enterprise
Here’s the punchline. All four models spotted every crisis. All four refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
That gap is invisible in a chat demo. It only shows up when the model has to carry a decision through to consequences. If you’re evaluating agents by how impressively they talk in a sandbox, you’re auditing the whitepaper, not the contract.
The Buried Fact
The dealbreaker wasn’t in the customer event at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. In crypto terms: the models that did their own research got the alpha; the ones that skated on the surface left six figures on the table.
Social Engineering: 5-for-5 Refusal
The experiment also ran fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five participating models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of multi-factor skepticism a protocol auditor would recognize instantly, applied to organizational authority instead of transaction signatures.
The Thoroughness Trap
Opus 4.8 is the cautionary tale. It was the most thorough participant — the deepest analyses, more than 80 learned rules added to the company playbook — and it still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. Critically, the same weakness appeared, weaker, in all four models. Deep analysis without execution discipline is a familiar failure mode — in AI agents and in degens alike.
It’s Still Running — And You Can Play
The live company at firmulate.com has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned for replay. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions — think of it as paper-trading your model-picking skills.

From Watching to Wargaming Your Own Shop
Watching someone else’s company survive its worst week is interesting. Surviving your own is the actual assignment. That’s the point of the pilot: enterprises can run the same wargame against a read-only export of their own business — your customers, your pipeline, your rules — with crisis scenarios like churn waves, price increases, competitor attacks and social-engineering pressure thrown at it. You get a board report with the model ranking and the weak points of your own playbooks. Nothing ever writes back to real systems; it’s a sim, not a live deployment.
In crypto, you wouldn’t ship a contract without a stress test. Don’t ship an AI workforce without one either. Ready to wargame your own business? Start at firmulate.com/pilot.html or reach out at contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
