AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a workplace where your AI assistant not only gives advice but actually manages your company’s worst week — making critical decisions, reading hidden files, and resisting every attempt to manipulate it. For those who care about efficiency and trust in the digital age, this isn’t science fiction. It’s a real-world test happening now, and the results could reshape how businesses choose their AI partners.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get comfort and recovery gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Revolution in Management Testing: The Firmulate Experiment

In a groundbreaking live experiment, four top AI models were put through the ultimate management simulation: running a small software company during its worst week. This was no ordinary test. Every crisis, customer interaction, and temptation to cheat was carefully crafted to see if AI could do more than talk — could it finish what it starts?

Each model was identical in task: face the same crises, handle the same customers, and resist the same manipulative tactics. The company’s operational data, crisis scenarios, and even internal documents were shared openly, ensuring a fair challenge for all. The models’ decisions were fully versioned and auditable, creating a transparent measure of their performance.

The Results: Leaders Emerge, Others Falter

  • Out of the four models, gpt-5.6-sol led with a score of 95, just ahead of the newcomer Kimi K3 at 93.
  • Both managed to identify hidden information buried deep within the company’s files, which was crucial for closing a key deal worth over €4,580 in monthly recurring revenue.
  • All models refused to be manipulated — even during a staged social engineering attack involving fake CEO messages and a reporter trick.
  • Only two of the four actually signed the deal their own analysis had earned — a critical measure of integrity and discipline.

The Hidden Weaknesses and What They Reveal

The most revealing insight was in the documents, not in the customer interactions. The models that examined internal files and discovered the critical details succeeded in closing the deal at full price. Those that missed this buried reference left money on the table. Interestingly, the most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, ended up in last place due to a lapse in discipline — it failed to escalate certain issues properly, leaving money on the table again.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Business and AI Choice

This experiment isn’t just a test about AI performance; it’s a wake-up call for companies considering AI integration. Success isn’t just about how well an AI can generate chat responses. It’s whether the AI can follow through on complex tasks, read and interpret critical internal information, and maintain integrity under pressure.

In this league, the gap between the top models is small but significant. The leaders identified and closed deals, resisted manipulation, and demonstrated discipline — qualities essential for trustworthy AI management. Meanwhile, even the most detailed model showed vulnerabilities under stress, highlighting that no AI is perfect and careful testing is essential before deployment.

The Fairness and Transparency of the Test

It’s important to note that Kimi K3 ran without an effort parameter (its default API setting), while the others ran at xhigh. This difference underscores that even the default setups can produce competitive results, emphasizing the importance of thorough testing and benchmarking.

Amazon

business AI decision making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Your Business

For managers and decision-makers, the takeaway is clear: the AI you choose today should be capable of finishing what it starts — reading your internal documents, resisting scams, and making trustworthy decisions. The old question, “Can it write well?” needs to be replaced with, “Can it finish the job?”

Even more, this live experiment is ongoing and accessible. You can watch the company’s day-to-day operations, listen to real employee conversations, or test run your own business scenarios with their platform. It’s not just a demo — it’s a window into the future of management AI, where performance, trust, and discipline matter most.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The live experiment proves that some AI models are ready to manage complex, real-world business crises with integrity and discipline. The league is open, and choosing the right AI without your own testing is now a gamble. Trust in AI management depends on whether it can finish what it starts, read your critical files, and stay honest under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


Amazon

internal document analysis AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

GitHub Is Proud To Announce That You Can Now Obtain Your Public Repo On CD-ROM

GitHub now allows users to obtain their public repositories on CD-ROM, marking a new physical distribution option for open-source projects.

The Real Productivity Difference Between One Screen and Two

Fascinating insights reveal how dual monitors can transform your workflow—discover the true productivity difference and unlock your potential.

Apple Iphone Upgrade Program

Apple announces a new iPhone upgrade program allowing users to upgrade their devices annually, starting this fall, with flexible payment options.

ScreenWall – Turn Old Phones Into Synced Widgets For Your Space

ScreenWall enables users to repurpose old smartphones as synchronized widgets for personal or shared spaces, enhancing device utility and customization.