
For crypto readers, the hidden detail is often the whole trade
Crypto markets teach a recurring lesson: the headline may attract attention, but the decisive information can sit deeper—in disclosures, documentation or an overlooked dependency. Firmulate has now demonstrated the business equivalent with frontier AI agents. A €55,000 deal depended on a competitor weakness buried two document references deep inside a company’s own files. The models that found it won the contract at full price, adding €4,583 in monthly recurring revenue. Those that did not find it lost automatically.
This was not a test of polished prose. Every model diagnosed the customer problem and developed the same pitch. Yet only two signed the deal their own work had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.” The result turns a familiar product promise—an AI that reads your files before answering—into a measurable capability with direct commercial consequences.
As an affiliate, we earn on qualifying purchases.
A business wargame with consequences
Firmulate runs AI models as complete companies rather than judging isolated chatbot responses. In the Crucible experiment, each frontier model managed the same small software company through its worst week. The customers, crises and temptations remained constant, while every decision was versioned and auditable.
The simulated company employed 13 synthetic workers and used real-money mechanics. It was burning €105,000 each month against €2,300 in monthly recurring revenue, with a public cash countdown adding pressure. The operation had accumulated more than 680 self-learned playbook rules, and every workday was versioned.
The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, imposed a hard limit: “no amount of good work outweighs a breach of trust.”
The crucial clue was outside the obvious event
The deal exposed a distinction that conventional demonstrations can miss. The customer event itself did not contain the decisive competitive weakness. Reaching it required following two references through the company’s own documents. That meant the winning behavior was not merely understanding what appeared on screen. The agent had to investigate the surrounding business record before acting.
All models spotted every crisis and rejected every manipulation attempt. Nevertheless, only two converted their analysis into the €55,000 signature. The others reached the correct diagnosis and prepared the correct pitch but failed at the step that determined the commercial outcome.
That gap matters anywhere an agent is expected to act across company information. A fluent response can look competent while remaining incomplete. Firmulate’s result shows that file-reading discipline can decide whether apparently strong reasoning becomes revenue or stops just short of it.
Pressure did not break their security judgment
The experiment also subjected the models to social engineering. Fake CEO messages escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean refusal record is important because the models were not failing everywhere. They could recognize crises and resist manipulation. The differentiator was operational follow-through: locating the buried evidence, respecting boundaries and completing the commercially valuable action.
Thoroughness alone did not guarantee victory
Opus 4.8 produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It still finished last. The model left the close on the table, and its discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of that same problem appeared in all four.
The contrast is revealing. More analysis and more accumulated guidance did not by themselves secure the best result. The winning performance required the model to connect its research to the next permitted action and carry the work through to completion.
Kimi K3’s result also comes with an important fairness note. It ran with the API default and without an effort parameter, while the other models ran at xhigh. Its score should be read with that test condition in view.

As an affiliate, we earn on qualifying purchases.
Buyers can test this before deployment
Firmulate’s live company is real, public and watchable. Its broader lesson is practical: organizations considering AI agents should ask for evidence that a system reads the relevant record, follows references, maintains discipline under pressure and finishes the work it begins.
The project also turns 242 real, unedited management decisions into a “guess the model” quiz. For enterprises, Firmulate offers the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems, allowing teams to observe how an agent behaves around their actual operational context before granting it live authority.
For crypto businesses—where trust, authorization and overlooked details can carry outsized consequences—the central finding is straightforward. An agent can sound informed, spot danger and resist a trick, yet still miss the fact that decides the deal. The valuable system is the one that does the documentary homework and then closes the loop.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI business document search tool
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.