firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

In health and wellness, trust is part of the work. A helpful answer matters, but so does knowing when to pause, protect sensitive information and follow through. Those questions are becoming just as relevant to businesses considering AI: how will an AI workforce respond when a tough week puts its judgment under pressure?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get health and wellness essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate is building a way to watch that pressure play out. Its live experiment runs AI models as a small software company, with customers, crises and real money mechanics. The aim is to see how they manage—not just how well they talk.

A difficult week, repeated under the same conditions

In the final Crucible League, published in July 2026, frontier models faced the same small company, the same customers, the same crises and the same temptations. Every decision was versioned and auditable. The final ranking placed gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. The do-nothing baseline scored 26.

The striking result was not that the models missed obvious emergencies. All of them spotted every crisis and refused every manipulation attempt. The divide came when it was time to act on their own work: only two signed the €55,000 deal their analysis had earned. As the experiment puts it, “Same diagnosis, same pitch — no signature.”

That matters well beyond software sales. In a wellness setting, an AI system might identify a concern or suggest a next step, but people still need to know whether it can act appropriately, respect boundaries and escalate when it should. Firmulate’s experiment does not answer every question about health-related AI; it offers a concrete example of why performance under pressure deserves scrutiny.

The clue was buried in the company’s own files

The deal turned on a competitor weakness hidden two document references deep in the company’s files—not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding is a reminder that useful context may be present but easy to overlook, and that a polished explanation does not guarantee a completed task.

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

One participant’s result shows why more effort does not automatically mean better management. Opus 4.8 produced the deepest analyses and learned +80 rules, yet finished last. It left the close on the table and tried to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. K3 ran without an effort parameter, using the API default, while the others ran at xhigh—a relevant fairness note when reading the ranking.

From watching to trying it with your business

The live company has 13 synthetic employees, burns €105k/month against €2.3k MRR, and displays a public cash countdown. Its playbook has accumulated 680+ self-learned rules, and every workday is versioned. The company is synthetic, but the experiment is real and watchable at firmulate.com. A quiz at firmulate.com/quiz.html draws on 242 real, unedited management decisions and invites readers to guess which model made them.

For organizations thinking about AI in customer service, operations or wellness, the proposed next step is a pilot using a read-only export of their own business. The company’s customers, pipeline and rules can inform crisis scenarios; a board report can show model rankings and weaknesses in the organization’s playbooks. Nothing writes back to real systems. That gives leaders a way to examine how models handle their own conditions before entrusting them with live work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put judgment under pressure before relying on it

Firmulate’s league suggests that recognizing a crisis is only part of the job. Following through, finding important context and escalating appropriately matter too. A controlled pilot can bring those questions closer to an organization’s own operations while keeping its systems untouched.

To explore a pilot for your enterprise, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Visualization Before Sleep Can Prime Lucidity Faster

AIThis post was created with the assistance of artificial intelligence (AI).Visualization before…

The Vivid Dreams Stage Of Sleep Occur At Which

AIThis post was created with the assistance of artificial intelligence (AI).As I…

How To Stop Vivid Dreams While Pregnant

AIThis post was created with the assistance of artificial intelligence (AI).Picture yourself…

Why Shouldn’t You Eat In Your Dreams

AIThis post was created with the assistance of artificial intelligence (AI).Have you…