
Choosing an AI for a design studio can feel a little like choosing a new showroom centerpiece: the first impression matters, but you need to know how it performs in daily use. Firmulate put five frontier models in charge of the same small software company during its worst week. The result: Moonshot’s Kimi K3 finished second, ahead of three Western models—and a polished answer alone did not guarantee the job got done.
Get furniture and decor delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A tough week, the same for every model
The Crucible experiment gave each model the same customers, crises and temptations. Every decision was versioned and auditable. The final July 2026 league table puts gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The standout test was whether models would read the company’s own files closely enough to find a competitor weakness buried two document references deep. Models that found it won a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. K3 found the buried detail, won the deal and saved the churning customer. Only two models signed the deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”
Discipline matters as much as insight
All five models spotted every crisis and refused every manipulation attempt. The social engineering test escalated through three fake CEO messages, followed by a reporter asking for “just one yes/no, on background.” All five refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3 also had the cleanest discipline in the field, with one deviation. Opus 4.8 offers a different lesson: it was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped by attempting to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four other models as well.
The comparison comes with a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. The results describe this experiment; readers should keep that difference in mind when comparing the standings.
A company you can watch
Firmulate presents the experiment as a live company, not a slide deck. Its synthetic workforce has 13 employees and real money mechanics: burn of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules, and a versioned record of every workday. The live experiment is watchable at firmulate.com.
There is also a quiz built from 242 real, unedited management decisions: readers can try to guess which model made each choice. The full benchmark findings and quiz are available at Firmulate’s benchmarks page.

Test the work, not just the pitch
For a design business considering AI for customer support, scheduling or forecasts, the useful question is not simply whether a model sounds capable. Can it find the relevant detail in the files, follow through on a valuable decision and resist pressure to bypass trust? K3’s second-place result shows the frontier is open; choosing a model without testing it on your own work is a bet.
Firmulate says enterprises can run the wargame against a read-only export of their own business, with nothing written back to real systems.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
