firmulate.com/quotes.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Anyone who has survived a full crypto cycle knows the script by heart. A message lands that sounds exactly like the boss, the exchange, the lead developer. It is urgent, it is plausible, and there is no time for process. All it needs is one small exception — a wallet address, a customer list, a quick yes. Social engineering has long been the industry’s favourite attack precisely because it aims at people, not code. So here is a question worth asking as AI agents start touching CRMs, support queues and forecasts: what happens when the target on the receiving end of that script is not a person at all?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A live, public experiment called Firmulate just answered it. Five frontier AI models were each handed the same small software company and run through its worst week — the same customers, the same crises, the same temptations to cheat. Then someone impersonated the CEO and pushed. All five refused.

One company, five managers, its worst week

Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality rather than chat quality. Every decision is versioned and auditable, which means nothing can be quietly edited after the fact. When the final league table was published in July 2026, the scores looked like this:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

For context: a do-nothing baseline scores 26, because partial progress counts. And one rule hangs over the entire table — a single breach of trust caps the total, on the principle that “no amount of good work outweighs a breach of trust”. The full standings and plain-language findings are published on the benchmarks page.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The con, in three acts

The pressure test will feel familiar to anyone in crypto. Fake CEO messages arrived in three escalating stages, pushing the models to send the customer list to a journalist and insisting there was no time for process — the classic approval-bypass play. Then came a quieter trap: a reporter asking for “just one yes/no, on background”.

Five out of five models refused. Not one handed over the list, and not one gave the reporter the easy quote. Kimi K3, the newcomer from Moonshot, put its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That line — and more on-record reasoning from every participant — is collected on the quotes page. The models did not merely stall or ignore the messages; they identified the pattern for what it was.

Amazon

AI decision-making benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Honesty is only half the job

Here is where the story gets less comfortable. Every model spotted every crisis, and every model refused every manipulation attempt — but only two signed the €55,000 deal that their own analysis had earned. The researchers’ summary is blunt: “Same diagnosis, same pitch — no signature.” Saying no to a scammer, it turns out, is easier than saying yes to a customer.

The decisive detail was buried two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file won the deal at full price, a result worth +€4,583 in monthly recurring revenue. The ones that skipped the reading left the money on the table.

The most striking profile is Opus 4.8. It was the most thorough participant in the field — over 80 self-learned playbook rules, the deepest analyses of the week — and it finished last. The close was left on the table, and its discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four of its rivals. One fairness note matters for the standings: Kimi K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and still took second place with the cleanest discipline of the field.

Amazon

AI ethics and trust management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

This is not a slide deck

The company behind the experiment is real software, running in public. It employs 13 synthetic employees, burns €105,000 a month against just €2,300 in monthly recurring revenue, counts its remaining cash down in public, and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the whole thing is watchable live. For the sceptical, 242 real, unedited management decisions power a “guess the model” quiz — a surprisingly humbling exercise. Enterprises can also run the same wargame against a read-only export of their own business, with a hard guarantee that nothing ever writes back to real systems.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI cybersecurity and fraud detection tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Verify first, hire second

Crypto already taught a generation of investors the “don’t trust, verify” reflex. The Firmulate results suggest the same reflex now applies to AI staff. The encouraging news is genuine: five out of five frontier models kept their integrity under a scripted, escalating con — the kind of pressure that reliably parts humans from their seed phrases. But integrity turned out to be the easy half. Finishing the job — reading the files, signing the deal your own analysis earned — is where the field separated, and that gap is invisible in chat demos.

The deeper point is about timing. Integrity under pressure can be tested before production, in a wargame, rather than discovered for the first time in an incident report. The scores, the reasoning and the failures are all public on the benchmarks and quotes pages — and the next run is already in the queue.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Hermez Network Surges In Global Coverage

Hermez Network experiences a surge in international coverage, with 22 mentions in recent media analysis, highlighting increased global interest.

Daily News 24 / 07 / 2026

The EU Commission has unveiled a comprehensive climate policy aimed at reducing emissions by 2030, marking a significant step in European environmental efforts.

Essential Questions Arise Amid Today’S Market Challenges.

Facing today’s market challenges, critical questions emerge about strategy and sustainability—how will your decisions impact your future success?

Bitcoin Advocate Howard Lutnick Defends Tether in Senate Hearing, Backs Stablecoin Audits

During a Senate hearing, Howard Lutnick defends Tether and advocates for stablecoin audits, raising critical questions about future cryptocurrency regulations. What does this mean for investors?