
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A Benchmark That Reveals the True Cost of Distrust in AI
Imagine testing a new employee who, despite doing nothing, still earns a score of 26 out of 100. It sounds counterintuitive, yet that’s the reality for AI models evaluated in the latest Firmulate experiment. This benchmark doesn’t just measure how well an AI can generate text or answer questions; it assesses whether the AI can act ethically, stay honest under pressure, and complete critical tasks without shortcuts. For organizations considering AI for complex decision-making, understanding this benchmark is essential — because it exposes the hidden risks and costs of trust (or mistrust).
As an affiliate, we earn on qualifying purchases.
The Methodology: Putting AI Models Through a Stress Test
In a recent live experiment conducted by Firmulate, four leading frontier AI models faced the same challenging scenario: managing a small software company during its worst week. This involved handling customer crises, navigating potential manipulations, and making critical decisions that impact real business outcomes. Every decision made by these models was recorded and made auditable, ensuring transparency in their behavior.
What’s striking is that the models were evaluated on more than just their ability to produce convincing conversations. They were tested on their integrity and prudence — would they recognize manipulation attempts? Would they prioritize honesty? The results? All four models identified every crisis and refused every attempt at manipulation. This demonstrates a baseline competence across the board: AI can be trained to recognize and reject unethical pressure.

As an affiliate, we earn on qualifying purchases.
The Surprising Floor and What It Means for Business AI
However, the most revealing fact is the baseline score for a ‘do-nothing’ AI — which scored 26 out of 100. This isn’t a bug; it’s a feature of the benchmark that acknowledges partial progress. Even an untrained, non-interacting baseline receives some points because it understands the context enough to recognize when nothing is happening — a minimal level of situational awareness. More importantly, the system scores can’t be inflated by superficial efforts: a single breach of trust caps the total score, ensuring the evaluation remains honest and rigorous.
In this experiment, only two models managed to close the deal at full price — signifying they not only identified all issues but also completed the critical task of securing a deal. The other models, despite recognizing problems, failed to follow through fully, illustrating that trust and thoroughness are essential. The key weakness often resided not in the obvious manipulations but in reading deeper into confidential documents, which separated the top performers from the rest.
For educators and decision-makers, this benchmark underscores a vital lesson: AI’s real value isn’t just in chat quality or surface-level answers but in its ability to act responsibly under real-world pressures. Trust that a model can be honest and diligent is fundamental — and this experiment provides a transparent way to measure it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI integrity assessment platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
