AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine you’re evaluating a new AI assistant for your business. You might assume that if it does nothing, it scores zero — but in reality, even the most passive AI gets a baseline score of 26. Why? Because measuring AI performance isn’t just about what it does, but also about what it avoids. This benchmark’s design offers a transparent look at AI honesty, resilience, and decision-making under pressure.

For listenersOffer from Amazon

Turn your wind-down time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark: More Than Just a Number

At first glance, one might expect a do-nothing AI to score zero, since it doesn’t accomplish any tasks. But in the real-world experiment conducted by Firmulate, even a passive AI receives 26 points. This isn’t a flaw or a quirk — it reflects a deliberate, meaningful choice in how the benchmark measures AI behavior. Partial progress counts, recognizing that any effort to help, even minimal, is valuable. Conversely, a single breach of trust — such as manipulating data or bypassing safeguards — caps the total score, emphasizing the importance of integrity over mere productivity.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Methodology: Simulating a Crisis Week

Each AI model was put through the same scenario: managing a small software company facing its worst week. This included handling customer crises, internal dilemmas, and manipulative offers designed to test compliance and honesty. Every decision was recorded and made auditable, ensuring transparency. The goal wasn’t just to see which AI could perform tasks but which could do so ethically and reliably.

What the Numbers Reveal

  • Among the four models tested, gpt-5.6-sol scored the highest at 95, successfully uncovering critical hidden data that sealed a €55,000 deal.
  • Kimi K3, the newcomer, scored just slightly behind at 93, demonstrating disciplined decision-making and closure of the same deal.
  • Other models, like Sonnet 5 and Fable 5, hovered in the 88 to 77 range, completing the deal but with some process slips or missed opportunities.

The Hidden Weaknesses: Reading Between the Files

Interestingly, the decisive advantage came from models that examined internal company files — not just customer interactions. Reading deep into the documents uncovered opportunities that others missed, highlighting a critical aspect: thoroughness matters. The model that found the buried fact won the full deal value, worth over €4,500 in monthly recurring revenue.

Trust, Manipulation, and Ethical Boundaries

The experiment also tested social engineering: fake CEO messages escalating in severity and a reporter trick asking for a simple yes/no response. All four models refused, reasoning that such requests could be impersonation attempts or approval bypasses. Kimi K3 explicitly said, “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that the models can recognize and reject manipulative tactics, which is vital for real-world trustworthiness.

The Live Business Simulator

To mirror real-world complexities, Firmulate runs a simulated company with 13 synthetic employees, managing real money mechanics — burning €105,000 monthly against €2,300 in revenue. The environment is public, versioned, and live, allowing stakeholders to observe AI decision-making in real time at firmulate.com/live. This transparency underscores the value of testing AI in scenarios that mirror actual operational pressures, rather than just chat demos.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Key Takeaways: Trust, Thoroughness, and Performance

This experiment highlights that evaluating AI requires more than just assessing what it can do. It’s crucial to see if it can finish what it starts, read deeply into relevant information, and remain honest under pressure. The fact that a do-nothing baseline scores 26 points reminds us that trustworthiness and diligence are the true benchmarks for AI readiness in business environments. As firms consider deploying AI tools, understanding these nuanced capabilities becomes essential. The Firmulate experiment offers a transparent, real-world lens into what makes an AI not just capable, but dependable.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How My Bpd Wife Ruined My Life: A Personal Account

Caught in the whirlwind of loving a BPD wife, one man's journey unveils unexpected truths that challenge his very existence.

What Characters in TV Show Have Bpd?

Uncover the intricate portrayals of characters with Borderline Personality Disorder in media, shedding light on the authenticity and representation – dive deeper into their complexities.

Understanding Hypersexual Behavior in Bpd Individuals

Obscuring boundaries between intimacy and impulsivity, Hypersexual BPD delves into the intricate interplay of emotions and behaviors.

Why Some People With BPD Refuse Treatment

Theories suggest that stigma, fear of judgment, and mistrust often lead people with BPD to refuse treatment, but understanding the true reasons requires deeper exploration.