📊 Full opportunity report: The Hidden AI Leaderboard That Starts After The Demo Wraps Up on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
A live experiment by Firmulate tests AI models in managing a small company during its worst week. Results show management skills, trust, and decision-making are critical, revealing a hidden leaderboard beyond traditional benchmarks.
In a groundbreaking live experiment, Firmulate has tested AI models in managing a small software company during its worst week, revealing a new dimension of AI evaluation focused on management quality, trust, and decision-making. The results, published after the July 2026 Crucible League, show that traditional benchmarks—like chat quality or coding scores—do not fully capture an AI’s effectiveness in real-world managerial tasks. For a deeper understanding, see the original analysis. This development matters because it shifts the focus toward how AI can responsibly and reliably handle complex, consequential business decisions.
The experiment involved five AI models competing to manage a simulated company facing multiple crises, including customer churn, PR issues, and financial pressures. Insights from this experiment are covered in the original analysis. Each model was evaluated on its ability to diagnose problems, communicate, negotiate, and execute decisions under strict trust constraints. The models were scored on a 100-point scale, with GPT-5.6-SOL leading at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Despite all models identifying crises and resisting manipulation attempts, only two successfully closed a critical €55,000 deal—highlighting that effective management involves more than just accurate diagnosis or eloquent responses.
One key finding was that models often failed to retrieve the most impactful information buried within the company’s files, leading to missed opportunities despite sounding informed. Additionally, models demonstrated strong resistance to social engineering tactics, refusing to escalate fake approval requests or disclose sensitive information, which is promising for trust and safety concerns. However, the experiment revealed that more comprehensive analysis and activity do not necessarily translate into better management outcomes, as seen with Opus 4.8, which added extensive rules but finished last due to discipline lapses in escalation processes.
The experiment’s design enforced a strict trust policy: any breach capped the overall score, emphasizing the importance of integrity over superficial performance. The live company used real cash flow mechanics, burning €105,000 monthly against €2,300 in monthly recurring revenue, making the stakes tangible. The setup included 13 synthetic employees and a detailed set of rules and decision logs, allowing observers to evaluate whether AI models prioritize the right actions, follow organizational context, and maintain honesty throughout complex scenarios. This approach is discussed in detail in the original analysis.
Why Management Skills Outperform Chat Benchmarks in AI Evaluation
This experiment underscores that AI’s ability to handle real management tasks—such as diagnosing crises, maintaining trust, and completing decisions—is more indicative of its readiness for enterprise deployment than traditional benchmarks like coding scores or conversational fluency. It highlights that effective AI management involves accountability, prioritization, and ethical boundaries, which are critical for responsible AI adoption in business environments. The findings suggest that future AI evaluations should incorporate scenario-based management tests to better predict real-world performance and risks.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Evolution of AI Benchmarks Beyond Traditional Tests
Until now, AI performance has largely been measured through coding leaderboards, chat arena ratings, and task-specific benchmarks. These tests focus on technical accuracy, language fluency, or problem-solving speed but neglect how models perform under pressure, in multi-faceted decision-making, or when managing organizational trust. The Firmulate experiment builds on recent efforts to evaluate AI in more realistic, consequential contexts, like crisis management and ethical decision-making, emphasizing that AI’s true value lies in its ability to manage complex, real-world scenarios.
Previously, the focus was on isolated capabilities; now, there is a shift toward integrated management performance, which considers how models handle conflicting priorities, incomplete information, and trust boundaries over extended periods. This approach aims to bridge the gap between theoretical benchmarks and practical enterprise use, where failure can have serious financial and reputational consequences.
“Traditional AI benchmarks measure isolated skills, but real management requires trust, prioritization, and accountability—dimensions that are only visible when models manage consequences over time.”
— Thorsten Meyer, lead researcher at Firmulate
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About Long-Term AI Management Performance
It remains unclear how these models will perform in live, fully integrated enterprise environments over longer periods. The experiment simulated a single week of crisis management, but real-world management involves ongoing adaptation, learning, and trust-building. Additionally, questions about scalability, model bias, and the ability to handle unforeseen crises are still open. The impact of different organizational contexts and the effectiveness of training or prompting strategies also require further investigation.
AI trust and safety assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps in AI Management Evaluation and Deployment
Future research will likely focus on extending these management simulations to longer timeframes, testing more diverse scenarios, and integrating models into actual business workflows. Companies interested in deploying AI for management tasks should consider scenario-based testing, focusing on trust, decision quality, and escalation protocols. Meanwhile, AI developers are expected to refine models to better handle organizational context, prioritize transparency, and adhere to ethical boundaries under real-world pressures. The ongoing development of benchmarks like Firmulate’s will help shape standards for responsible AI management capabilities.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main purpose of the Firmulate experiment?
The experiment aims to evaluate AI models’ ability to manage a real business scenario, focusing on decision-making, trust, and handling crises, beyond traditional chat or coding benchmarks.
All tested models successfully refused manipulative requests such as fake approval or impersonation attempts, demonstrating strength in trust boundaries.
What are the key limitations of this experiment?
The simulation covers only a single week, and it remains uncertain how models will perform over longer periods or in fully operational environments with unpredictable crises.
Why is management performance more important than chat quality?
Management involves accountability, trust, and effective decision-making that directly impact business outcomes, which chat quality alone cannot measure.
What should companies consider before deploying AI for management tasks?
Organizations should assess whether models can read organizational context, escalate issues properly, and maintain honesty under pressure, not just generate convincing responses.
Source: ThorstenMeyerAI.com