
Get movie nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Could AI models be the new front-runners in business decision-making?
In a real-world experiment that pits AI models against each other in a high-stakes business simulation, a newcomer has stunned industry watchers by finishing just behind the top contender. This isn’t science fiction; it’s an ongoing live test at firmulate.com, where AI is challenged to run a small software company through its worst week.
As an affiliate, we earn on qualifying purchases.
The Experiment: Testing AI in the Trenches
Four frontier AI models were tasked with managing a realistic company facing common crises: customer churn, security threats, and ethical dilemmas. Each model operated in identical conditions, with decisions fully documented and auditable for transparency. The goal was to see which AI could not only identify problems but also act decisively and honestly under pressure.
The models included the well-known gpt-5.6-sol, the recent entrant Moonshot’s Kimi K3, and two versions of Sonnet. The experiment aimed to evaluate core management qualities—trustworthiness, problem-solving, and discipline—beyond just chat skills.
Key Findings: The Results Speak Volumes
- All four models detected every crisis and refused manipulative attempts, demonstrating strong ethics and vigilance.
- Only two models achieved the critical goal: closing a €55,000 deal that was earned through accurate diagnosis and truthful pitching.
- Kimi K3, the newcomer from Moonshot, scored a near-perfect 93—just behind the leader at 95—and managed to uncover a hidden, crucial detail in the company’s files that sealed the deal.
- Interestingly, the top scorer, gpt-5.6-sol, identified the buried fact and closed at full price, showing the potential for high-level strategic insights.
The Hidden Weakness and the Discipline of Honesty
The experiment revealed a subtle but critical flaw: deeper document analysis was the key to winning. All models that read beyond surface data succeeded in securing the deal, emphasizing the importance of thorough information processing. The most disciplined participant, Kimi K3, maintained integrity and refused to deviate even under social engineering attempts—fake CEO messages and reporter tricks.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Business and AI?
In a world where AI tools increasingly influence customer relations, operations, and strategy, this experiment underscores a vital point: the quality of AI decision-making isn’t just about generating convincing chat. It’s about consistency, honesty, and the ability to finish what they start—even in high-pressure situations. The experiment’s real company is live at firmulate.com, running every business day, with real money mechanics and a burn rate of €105k/month against €2.3k MRR.
Why Should You Care?
For executives and managers, the takeaway is clear: choosing an AI model based solely on superficial performance is a gamble. The real test lies in how well these models handle complex, messy, and ethically fraught scenarios. The league table, with scores from 73 to 95, shows the landscape is open for newcomers—like Kimi K3—to challenge established players, with a clear winner emerging only after rigorous real-world tests.
AI ethics and transparency tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fairness and Transparency in AI Evaluation
It’s important to note that Kimi K3 ran without an effort parameter (the API’s default setting), while the other models ran at xhigh. This difference didn’t prevent K3 from excelling, indicating robustness in its decision-making process.
As an affiliate, we earn on qualifying purchases.
Final Word: The Future of AI in Business
The experiment at firmulate.com proves that AI can be more than just a chatty assistant. It can act as a disciplined, trustworthy partner capable of managing real business crises and securing deals. As the league grows more competitive, the best choice will depend on thorough testing—just like in this live, ongoing experiment. Before you hire your next AI, consider running your own ‘wargame’ to see how it performs under pressure.

Key takeaway:
In high-stakes business scenarios, the true value of an AI model lies in its ability to read deeply, stay honest, and finish what it starts. Kimi K3’s strong performance shows that newcomers can beat established players when tested rigorously—making AI selection a strategic decision, not just a quick choice.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
