📊 Full opportunity report: Unveiling AI’s True Working Style With A Simple Management Test on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
A live management test compares AI models’ ability to handle a simulated company’s worst week. The results highlight differences in diligence, trust, and decision execution, offering new insights into AI management and decision-making.
Five AI management models were tested in a live, real-time simulation of a company’s worst week, revealing significant differences in their ability to follow through on decisions, protect trust, and complete critical tasks. This experiment is the first of its kind to directly compare AI decision-making in a complex business environment, making it highly relevant for enterprises evaluating AI management capabilities.
The experiment, conducted by Firmulate.com, involved five frontier AI models managing a small software company facing crises, customer issues, and operational pressures. For more on AI decision testing, see the original analysis. Each model was tasked with handling 242 real, unedited management decisions, including crisis response, sales negotiations, and security protocols. The models’ performance was scored based on their ability to diagnose problems, take decisive action, and maintain trustworthiness. The top performer, gpt-5.6-sol, scored 95 points out of 100, while others lagged behind, with the baseline scoring just 26.
Despite all models recognizing crises and resisting manipulation attempts, only two successfully signed a crucial €55,000 deal, demonstrating the gap between analysis and execution. Notably, the model Opus 4.8, despite producing the most thorough analysis, failed to complete key operational steps, such as escalating issues properly, which impacted its overall score. The experiment underscores that effective management relies not just on understanding but on decisive action and follow-through.
Implications for AI Management and Business Automation
This experiment demonstrates that AI models vary significantly in their ability to translate diagnosis into action, a critical factor for enterprise automation. It highlights the importance of testing AI decision-making in realistic scenarios before deploying them operationally. The results suggest that AI’s value in business lies not only in analysis but in reliably executing decisions that impact revenue and trust, emphasizing the need for careful evaluation of AI management personalities.
AI decision-making management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of AI Decision-Making Tests in Business
Previous AI demonstrations often focused on isolated capabilities like language understanding or problem diagnosis. However, real-world management requires integrating analysis with execution, trust, and security. The Firmulate experiment builds on ongoing efforts to evaluate AI in operational roles, using a live, unfiltered company simulation to observe how models perform under pressure. The July 2026 league results follow earlier benchmarks that showed AI models can recognize crises but struggle with decisive follow-through.
“Testing AI models against real management tasks reveals their true operational strengths and weaknesses, beyond superficial analysis.”
— Firmulate.com
As an affiliate, we earn on qualifying purchases.
What Aspects of AI Performance Are Still Unclear
It remains unclear how these models would perform in different industry contexts or with more complex, longer-term management tasks. The experiment focused on a specific scenario—managing a software company’s worst week—and results may not generalize to other operational environments. Additionally, the impact of different operational parameters, such as API settings or training data, on performance is still being studied.
As an affiliate, we earn on qualifying purchases.
Future Steps for Testing and Deploying AI Management Models
Researchers plan to expand testing to different industries and longer management cycles, aiming to identify the best practices for AI deployment in real-world business settings. Enterprises are encouraged to run their own simulations, using similar live tests to assess AI models’ readiness before granting operational authority. Further benchmarking will refine understanding of how AI personalities influence decision quality and follow-through.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why is this live management test significant?
This test provides a realistic assessment of AI models’ ability to handle complex management tasks, revealing strengths and weaknesses that are often hidden in traditional benchmarks.
What does the experiment reveal about AI decision-making?
It shows that recognizing crises is not enough; AI must also execute decisions reliably. The ability to follow through and complete critical actions distinguishes top-performing models.
Can these AI models be trusted for operational management?
The experiment indicates that while models can identify problems, their operational discipline varies. Careful testing and validation are necessary before deployment in live business environments.
What are the limitations of this experiment?
The scenario is specific to managing a software company’s crisis week, and results may differ in other sectors or with different operational complexities. Further research is needed to generalize findings.
How can companies use this testing approach?
Businesses can replicate similar live simulations with their own data to evaluate AI models’ real-world readiness before operational use, ensuring better alignment with their specific needs.
Source: ThorstenMeyerAI.com