
Imagine a business that operates entirely without human staff, facing daily crises, tough decisions, and ethical dilemmas—all in front of your eyes. This isn’t fiction; it’s the live experiment by Firmulate, where artificial intelligence models run a small software company through its worst week, revealing how AI can mimic human decision-making under pressure—and where it still struggles.
The Live Experiment: AI as a Company Leader
At the heart of this experiment is a small, virtual company managed entirely by AI models, each acting as a synthetic employee. These models are tested against a series of real-world crises—customer complaints, financial dilemmas, and ethical challenges—on the company’s 183rd day of operation. Every decision is versioned and made auditable, creating a transparent window into AI behavior in high-stakes scenarios.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the AI Models Are Tested
The models are pitted against each other in a rigorous ‘wargame.’ Four frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—each run the same simulated week of a business crisis. They encounter the same customers, face the same operational setbacks, and are subjected to manipulative tactics intended to test their integrity.
Each model’s performance is scored based on their ability to identify crises, resist manipulation, and execute profitable deals. For instance, while all models detected every crisis and refused unethical manipulation attempts, only two managed to close the €55,000 deal that their analysis justified. The other two either left the deal on the table or failed to follow through, revealing critical weaknesses in discipline and decision execution.

Key Lessons from the Experiment
- Integrity under pressure is crucial: All models refused manipulative tactics, illustrating that AI can be designed to prioritize honesty, even when incentivized otherwise.
- Understanding the buried facts matters: The decisive advantage came from reading deeper into internal documents—something that models which explored these references won at full price (+€4,583 MRR).
- Discipline influences success: The most thorough model, Opus 4.8, with over 80 learned rules, failed to close a deal due to lapses in discipline, highlighting that extensive training alone isn’t enough without consistent execution.
- Transparency and versioning are vital: Every decision, whether successful or not, is logged and auditable, showing that building AI decision-making processes with transparency is possible—and necessary for trustworthiness.
This experiment isn’t just about AI in business; it touches on core human concerns: trust, discipline, and the ability to act ethically when stakes are high. The real-world implications extend to how AI might soon influence customer relations, support, and decision-making in your own organization.
Why It Matters for Mental Health & Psychology
In a way, observing AI navigating these crises mirrors human dilemmas—how we respond to pressure, temptation, and ethical quandaries. The experiment underscores the importance of integrity, discipline, and the capacity to face uncomfortable truths—concepts central to psychology and mental health. Just as AI models can be trained to refuse manipulation, humans benefit from developing resilience and ethical clarity in stressful situations.
As this transparency-driven experiment continues, it offers a mirror to our own decision-making processes, encouraging reflection on how we can foster honesty and discipline in ourselves—and how technology might support us in doing so.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html