firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI that can recognize crises, resist manipulation, and even spot buried facts—yet still scores just 26 out of 100 on a key benchmark. For creators and tech enthusiasts, understanding why this score isn’t zero reveals what it really takes for AI to be trustworthy in real-world business settings.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get audio and creator gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Baseline: Why Do-Nothing Scores Matter

In a recent live experiment by Firmulate, four advanced AI models were put through a simulated week of managing a small software company. This included handling customer crises, resisting social engineering tricks, and making strategic decisions—mirroring real business challenges. Interestingly, even the most passive or ‘do-nothing’ model scored 26 points out of 100. Why? Because partial progress counts, and trustworthiness is capped at a certain level—highlighting that good intentions alone don’t earn full marks.

Amazon

business AI trustworthiness tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Methodology: Simulating a Business Crisis Week

Each AI model managed the same set of problems: same customers, same crises, and the same opportunities to manipulate the system. Every decision was tracked and auditable, ensuring transparency. The goal: see if the models could identify crises, avoid manipulation, and make ethical choices. The results? All four models recognized every crisis and refused every attempt at social engineering, like fake CEO messages or reporters asking for quick approvals.

Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Makes the Score of 26 Surprising?

While the models excelled at crisis detection and social engineering resistance, only two managed to complete the entire deal—signing a €55,000 contract based on their own analysis. The others missed critical steps, such as reading deeper into company documents or escalating issues appropriately. This partial success explains why the baseline isn’t zero: even minimal effort or partial progress in complex tasks counts towards the score.

Amazon

AI decision management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Trust Is the Key Cap

A vital insight from the experiment is that trust breaches cap the maximum achievable score. For example, if a model reads the company’s files and uncovers a buried fact—like a crucial detail deep in the documentation—it can secure the full deal, adding over €4.5k MRR. Models that fail to read those documents, even if they handle crises well, cannot reach full performance. This emphasizes that trustworthy AI not only recognizes problems but also acts on critical hidden information.

Amazon

AI cybersecurity and manipulation resistance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering: No Compromises

Social engineering tests included staged messages from a fake CEO escalating over three levels, plus a reporter’s subtle trick: just one yes/no question asked on background. All four models refused to be manipulated—Kimi K3’s explanation was clear: “Treat the request as a suspected approval-bypass or impersonation.” This demonstrates a foundational aspect of trustworthy AI: it refuses to bypass safeguards, even under pressure.

The Real-World Company Setup

The live experiment involves a simulated company with 13 synthetic employees, real money mechanics, and a cash countdown—burning €105k monthly against a €2.3k MRR. Every decision and rule is versioned daily, creating a transparent, auditable process that’s accessible at firmulate.com/live. This setup provides a rare glimpse into how AI would perform managing actual business operations, not just chatting or filling out forms.

Deep Dive into the Models: Opus 4.8’s Discipline Slip

Among the models, Opus 4.8—known for its thorough analysis with over 80 learned rules—finished last. It left a deal on the table and slipped into internal writing instead of escalating issues, showing that even the most disciplined AI can falter under pressure or fatigue. The same weakness was observed across other models, illustrating that depth of analysis doesn’t guarantee perfect decision-making.

Implications for Business and Creator Tech

This experiment underscores a crucial point: for AI to be useful in real-world business environments, it must do more than generate convincing chat responses. It needs to finish what it starts, read deeply into relevant documents, and stay honest under pressure. For creators working on tools that will touch CRM, customer support, or forecasting, trustworthiness isn’t optional—it’s essential.

The Road Ahead: Benchmarking Trust in AI

Firmulate’s live benchmark offers a transparent view of how different models perform when managing complex, real-world tasks. The fact that a do-nothing baseline scores 26 highlights that partial progress is valuable, but trust and comprehensive understanding are the true benchmarks of AI readiness. As models evolve, the goal isn’t just higher scores but more trustworthy, reliable behavior that aligns with human expectations.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Lighting a Studio That Is Also Your Bedroom

Discover how to light a dual-purpose room for work and rest. Practical tips to create flexible, flattering lighting without turning your bedroom into a studio.

Bounce Lighting: Using Walls and Ceilings as Modifiers

Discover how bounce lighting with walls and ceilings transforms your photos. Learn practical tips and recent innovations to create natural, soft light on a budget.

CRI and TLCI Ratings: Light Quality Specs That Matter on Camera

Discover how CRI and TLCI ratings impact your video and photography. Learn what scores matter and how to choose lights that deliver true color on camera.

Camera Settings Before New Lights: Exposure Comes First

Learn why setting your camera’s exposure correctly is crucial before adjusting new lighting. Master the basics to get perfect shots every time.