firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Crypto businesses already know that a polished pitch is not the same as a system that holds up under pressure. Firmulate’s live experiment puts AI models through a company’s worst week, testing whether they can read the evidence, protect trust and follow through when money is on the line.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A test beyond the demo

Firmulate ran frontier models through the same small software company and the same set of customers, crises and temptations. Decisions were versioned and auditable. The experiment is presented as a live company, with synthetic employees and real money mechanics; readers can watch it at Firmulate.

The final Crucible league, dated July 2026, puts gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. Kimi’s result makes the field look more open: it finished ahead of three of the four Western models in the test.

Amazon

AI cybersecurity testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading the files mattered

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The decisive weakness in a competitor’s position was buried two document references deep in the company’s files, rather than in the customer event. Models that found it won the deal at full price, worth +€4,583 in monthly recurring revenue.

Kimi K3 found that buried fact, won the deal, saved a customer who was considering leaving and resisted all three baits. It had one deviation, the fewest in the field. The results point to a practical distinction for businesses considering AI agents: noticing a problem is different from completing the work that follows.

The manipulation test included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. Kimi’s on-record reasoning described the request as a “suspected approval-bypass / possible impersonation.”

Amazon

AI decision-making simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness did not guarantee a finish

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it placed last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

The company behind the test has 13 synthetic employees, burns €105k a month against €2.3k in monthly recurring revenue, and displays a public cash countdown. Its playbook has more than 680 self-learned rules, and each workday is versioned. Firmulate also offers a quiz built from 242 real, unedited management decisions, asking visitors to guess which model made each choice.

Amazon

crypto business risk assessment AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Run your own test

For crypto firms weighing AI for customer support, operations or forecasting, the standings are a reason to test models against the work they would actually do. Firmulate says enterprises can run the same wargame against a read-only export of their business; nothing writes back to real systems. The benchmark page publishes the results and findings.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI model testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The takeaway

Kimi K3’s second-place result shows that the leaders are not a foregone conclusion. But a league table cannot tell a crypto business which model will handle its own customers, files and approval risks. Choosing without running a relevant test is a bet.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Crypto Spot Trading Hits New Heights, Fueling $11.3T Market Activity

Get ready to explore the explosive rise of crypto spot trading, as market activity soars to $11.3 trillion—what’s driving this unprecedented growth?

IRS Refund Delays: What to Know Before Filing This Year

Can your tax refund be delayed this year? Discover essential tips to ensure a smoother filing process and avoid common mistakes.

Will Darline Graham Nordone Be The New Republican Nominee For Senate In South Carolina?

Speculation grows around Darline Graham Nordone’s potential to secure the Republican nomination for South Carolina Senate seat in upcoming elections.

Securitize and tZERO clash over patents as race to bring Wall Street onchain heats up

Securitize and tZERO clash over patent rights amid growing efforts to digitize Wall Street assets, raising questions about industry consolidation and innovation.