AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A furniture showroom has its own pressure points: a late delivery, a customer ready to walk, a rival offering a sharper price, and a manager juggling messages that may not be what they seem. An AI assistant can sound convincing when asked for a product description. The tougher question is whether it can make sound decisions when the whole business is under strain.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get furniture and decor delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That is the question behind Firmulate, a public experiment that puts AI models in charge of the same small software company and lets visitors watch the results. Its next step is aimed at businesses that want to see how their own playbooks hold up.

A business has to be tested in context

In the final Crucible League, published in July 2026, each frontier model faced the same company, customers, crises and temptations. The experiment follows decisions through a difficult week, rather than judging an isolated conversation. Each decision is versioned and auditable.

The results expose a gap that a polished answer can hide. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s summary is blunt: “Same diagnosis, same pitch — no signature.”

One explanation lay deep in the company’s own documents. The decisive weakness in a competitor’s position was buried two document references into the files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding points to a practical challenge for any business: useful information can be present and still go unnoticed when decisions depend on connecting details across company records.

Trust and follow-through matter as much as insight

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

But refusing manipulation was not the only measure of performance. In the final league, gpt-5.6-sol placed first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. The league’s integrity rule is uncompromising: “no amount of good work outweighs a breach of trust.”

Opus 4.8 makes the difference between diligence and results especially clear. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There is a fairness caveat for readers interpreting the ranking: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

A watchable company, then a business-specific pilot

Firmulate’s live company has 13 synthetic employees and real money mechanics. Its burn is €105k per month against €2.3k MRR, with a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned. The experiment is real and watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.

For a company considering AI agents in areas such as customer service, sales or forecasting, the experiment offers a route from observing a live company to testing decisions against its own business context. Firmulate’s enterprise pilot uses a read-only export to create a digital twin, then runs crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

See how your own playbooks respond

A benchmark can show how models perform in a shared scenario. A pilot can put your company’s own information and pressure points into view, while keeping the exercise read-only. To discuss a Firmulate pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Midnight Sun Planting Calendars

Inevitably, mastering midnight sun planting calendars can transform your garden—discover how to optimize growth under continuous daylight.

Procedural SVG Generation: A Look Inside “Rosarium Atelier | Cathedral Rose Window Restoration” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“Rosarium Atelier…

AI Models Stand Firm Against Social Engineering—A New Benchmark in Trust and Integrity

In a live experiment, five top AI models faced social engineering tricks and all refused manipulation, proving trustworthiness can be tested and verified before deployment.

The Daily Chore Rhythms That Changed With the Midnight Sun

Just how do daily chores and routines transform under the midnight sun, and what surprising effects does this shift have on everyday life?