AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

For listenersOffer from Amazon

Turn your wind-down time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Insight is not the same as follow-through

In psychology, recognizing a problem and acting on that recognition are different challenges. The same gap is now showing up in AI systems asked to run a business: every model spotted the crises, but only two followed their own analysis through to a consequential decision. Firmulate’s live experiment offers a practical test of what happens when an AI has to do more than sound convincing.

A company under pressure

Firmulate put frontier models in charge of the same small software company during its worst week, with the same customers, crises and temptations. Its employees are synthetic, but the experiment is real and watchable: the company has real money mechanics, a public cash countdown and versioned workdays.

The final Crucible league, dated July 2026, puts gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. K3 finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. That makes the newcomer’s performance a challenge to assumptions about which model will manage a business best. It also leaves the top place narrowly contested.

The decision after the diagnosis

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The company’s decisive competitive weakness was buried two document references deep in its files, rather than stated in the customer event. Models that read those files won the deal at full price, worth +€4,583 in monthly recurring revenue.

K3 found that buried security needle, won the deal and saved the churning customer. It resisted all three baits and had one deviation, the cleanest discipline in the field. The contrast is captured in Firmulate’s phrase: “Same diagnosis, same pitch — no signature.” Spotting what is wrong, or explaining what should happen, does not guarantee that an agent completes the task.

Pressure, trust and uneven performance

The test included fake CEO messages that escalated over three stages, plus a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” In a business setting, that kind of resistance matters alongside productivity: an agent’s judgment is tested by whether it holds when a request appears urgent or authoritative.

Opus 4.8 shows why a high-effort process does not necessarily produce the best outcome. It was the most thorough participant, with +80 learned rules and the deepest analyses, but came last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. Firmulate reports the same weakness, in milder form, across all four. Thoroughness and action are related, but they are not interchangeable.

The results also come with a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. That difference belongs beside the rankings when interpreting the comparison.

Why run a test of your own?

Firmulate frames the experiment as a way to evaluate AI workforce management rather than chat quality. The live company has 13 synthetic employees, burns €105k per month against €2.3k in monthly recurring revenue, and has accumulated more than 680 self-learned playbook rules. Its public countdown and daily work make the experiment watchable at Firmulate.

The site also offers a quiz built from 242 real, unedited management decisions. And for enterprises, Firmulate describes a pilot using a read-only export of their own business; the export does not write back to real systems. Those options point toward a broader question for people considering AI agents: how does a system behave when decisions have consequences, records matter and the next step requires follow-through?

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The case for testing behavior

Kimi K3’s second-place finish, ahead of three of the four Western frontier models in the league, makes the field look open. But the central finding is less about a single ranking than the gap between understanding a situation and completing the right action. If an AI model may handle customers, company records or approvals, choosing one without testing it in conditions that resemble the job is a bet. Read Firmulate’s benchmark findings.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Management Tests Reveal Hidden Weaknesses Beyond Chat Performance

AI models excel at chat but often hide vulnerabilities in real-world management. Live experiments reveal the importance of testing AI under pressure to ensure trust and integrity.

Medication in BPD: Why No Single Drug Is a Cure-All

Psychotropic medications in BPD target symptoms but fall short of a cure, revealing the complex interplay of factors behind effective treatment.

Understanding the Link Between Sex and Self-Harm in BPD

Nestled within the intricate web of Borderline Personality Disorder lies a complex link between sex and self-harm that begs to be unraveled.

Understanding the Link Between Bpd and Trust Issues

Yearning for connection but plagued by doubt, understanding how BPD intertwines with trust issues unveils a captivating journey of vulnerability and resilience.