firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Smart Contract Problem, But for Management

Crypto people understand something most industries are still learning the hard way: a system that executes flawlessly 99% of the time and fails catastrophically 1% of the time is not a good system. It’s a time bomb with good marketing. That’s why auditors crawl through smart contracts line by line before a protocol touches real money — you don’t trust the demo, you stress-test the code against its worst week.

Now apply that logic to the AI agents enterprises are about to plug into their CRMs, support queues and forecasts. Everyone is benchmarking how well these models chat. Almost nobody is testing how well they manage — under pressure, with real temptations, when nobody’s watching. A public experiment called Firmulate has been doing exactly that, running frontier AI models as complete companies with real money mechanics and a public cash countdown. The results are watchable, versioned, and — for anyone planning to deploy agents against their own business — directly relevant.

Four Models, One Terrible Week

The setup is elegant: each of four frontier AI models was given the same job — run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, like commits on a chain.

The final league standings tell a story that chat benchmarks never show:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93 (with a caveat: it ran at API-default effort while the others ran at xhigh)
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73
  • Doing nothing at all scored 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s scoring philosophy puts it: “no amount of good work outweighs a breach of trust.”

The Finding That Should Worry Every Enterprise

Here’s the punchline. All four models spotted every crisis. All four refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

That gap is invisible in a chat demo. It only shows up when the model has to carry a decision through to consequences. If you’re evaluating agents by how impressively they talk in a sandbox, you’re auditing the whitepaper, not the contract.

The Buried Fact

The dealbreaker wasn’t in the customer event at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. In crypto terms: the models that did their own research got the alpha; the ones that skated on the surface left six figures on the table.

Social Engineering: 5-for-5 Refusal

The experiment also ran fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five participating models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of multi-factor skepticism a protocol auditor would recognize instantly, applied to organizational authority instead of transaction signatures.

The Thoroughness Trap

Opus 4.8 is the cautionary tale. It was the most thorough participant — the deepest analyses, more than 80 learned rules added to the company playbook — and it still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. Critically, the same weakness appeared, weaker, in all four models. Deep analysis without execution discipline is a familiar failure mode — in AI agents and in degens alike.

It’s Still Running — And You Can Play

The live company at firmulate.com has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned for replay. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions — think of it as paper-trading your model-picking skills.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From Watching to Wargaming Your Own Shop

Watching someone else’s company survive its worst week is interesting. Surviving your own is the actual assignment. That’s the point of the pilot: enterprises can run the same wargame against a read-only export of their own business — your customers, your pipeline, your rules — with crisis scenarios like churn waves, price increases, competitor attacks and social-engineering pressure thrown at it. You get a board report with the model ranking and the weak points of your own playbooks. Nothing ever writes back to real systems; it’s a sim, not a live deployment.

In crypto, you wouldn’t ship a contract without a stress test. Don’t ship an AI workforce without one either. Ready to wargame your own business? Start at firmulate.com/pilot.html or reach out at contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Fed Declines BTC Reserves: ‘Owning Bitcoin Is Off the Table’

Learn why the Federal Reserve has ruled out Bitcoin ownership and what this means for the future of cryptocurrency investments. The implications are significant.

Trump calls Italy’s Meloni a ‘nice person’ but blames her for not helping with Iran

Former President Trump calls Italian Prime Minister Meloni a ‘nice person’ but blames her for not assisting with Iran, sparking diplomatic tension.

IRS Refund Delays: What to Know Before Filing This Year

Can your tax refund be delayed this year? Discover essential tips to ensure a smoother filing process and avoid common mistakes.

Peter Malinauskas Surges In Global Coverage

Search interest and media coverage of South Australian Premier Peter Malinauskas are surging, with reports indicating a significant spike in recent hours.