firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Choosing an AI for a design studio can feel a little like choosing a new showroom centerpiece: the first impression matters, but you need to know how it performs in daily use. Firmulate put five frontier models in charge of the same small software company during its worst week. The result: Moonshot’s Kimi K3 finished second, ahead of three Western models—and a polished answer alone did not guarantee the job got done.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get furniture and decor delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A tough week, the same for every model

The Crucible experiment gave each model the same customers, crises and temptations. Every decision was versioned and auditable. The final July 2026 league table puts gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The standout test was whether models would read the company’s own files closely enough to find a competitor weakness buried two document references deep. Models that found it won a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. K3 found the buried detail, won the deal and saved the churning customer. Only two models signed the deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

Discipline matters as much as insight

All five models spotted every crisis and refused every manipulation attempt. The social engineering test escalated through three fake CEO messages, followed by a reporter asking for “just one yes/no, on background.” All five refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3 also had the cleanest discipline in the field, with one deviation. Opus 4.8 offers a different lesson: it was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped by attempting to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four other models as well.

The comparison comes with a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. The results describe this experiment; readers should keep that difference in mind when comparing the standings.

A company you can watch

Firmulate presents the experiment as a live company, not a slide deck. Its synthetic workforce has 13 employees and real money mechanics: burn of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules, and a versioned record of every workday. The live experiment is watchable at firmulate.com.

There is also a quiz built from 242 real, unedited management decisions: readers can try to guess which model made each choice. The full benchmark findings and quiz are available at Firmulate’s benchmarks page.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the work, not just the pitch

For a design business considering AI for customer support, scheduling or forecasts, the useful question is not simply whether a model sounds capable. Can it find the relevant detail in the files, follow through on a valuable decision and resist pressure to bypass trust? K3’s second-place result shows the frontier is open; choosing a model without testing it on your own work is a bet.

Firmulate says enterprises can run the wargame against a read-only export of their own business, with nothing written back to real systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Is the Difference Between a Smart Switch and a Smart Relay?

Discover the pivotal differences between smart switches and smart relays, and find out which one perfectly suits your home automation needs.

No Learn Button on Garage Door Opener? Here’s What to Do

You might think a garage door opener without a learn button is hopeless, but there are effective solutions waiting for you to discover.

Portable Lifestyle Displays: Hisense RoamView Brings Big-Screen Viewing Anywhere…

Hisense introduces RoamView, a portable display designed for mobile, big-screen viewing, marking a new trend in portable lifestyle technology with rising interest.

Smart Bulb Vs Smart Switch: Decide Today

Compare smart bulbs and switches to find the perfect lighting solution for your home, but understanding their differences is essential before choosing.