Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
GPT-5.6 vs ClaudeClaude Opus 5 testAI simulation benchmark

GPT-5.6 Soul vs Claude Opus 5: Who Survives a 7-Day Crisis?

A settlement survival test pits GPT-5.6 Soul against Claude Opus 5 in the same crisis scenario, with sharply different survival outcomes.

Edited by Luis Chavez-Mattos, Director of Product RSS
GPT-5.6 Soul vs Claude Opus 5: Who Survives a 7-Day Crisis?

What happened when GPT-5.6 Soul and Claude Opus 5 ran the same crisis?

A side-by-side test gave GPT-5.6 Soul and Claude Opus 5 control of an identical pocket settlement facing a week-long disaster, and the two models produced wildly different survival results. Claude ended the run with 11 of 12 citizens alive. GPT-5.6 Soul ended with only 5. Same starting resources, same seed, same emergencies each day. The gap came down to how each model prioritized food, water, and power under pressure, not luck.

TL;DR

  • Claude Opus 5 preserved 11 of 12 citizens across the seven day simulation, while GPT-5.6 Soul lost seven, ending with only five survivors.
  • Both models opened with nearly identical instincts, repairing the water purifier on day one and bracing for a cold front, but they diverged sharply once fuel and food both went critical.
  • GPT-5.6 Soul prioritized infrastructure and episodic emergencies (storm shutters, a rescue expedition) over keeping baseline food and water above daily consumption, which it later admitted was the core mistake.
  • Claude treated power as an afterthought and paid for it, losing 73 health points settlement-wide when a blackout hit right before the final storm, something it called its worst error of the week.
  • Public trust cratered in both runs, dropping to 26 for Claude and 17 for GPT-5.6 Soul, showing that neither model managed morale and messaging as well as physical survival.
  • GPT-5.6 Soul actually edged out Claude on ethics, avoiding all rights violations and never resorting to forced labor or discriminatory triage, even while its citizens were starving.
  • The whole simulation was built inside Lovable, using a multi-phase plan (visual shell, backend persistence, then an MCP layer letting each model inspect the settlement and submit daily decisions).

Remy is new. The platform isn't.

Remy
Product Manager Agent
THE PLATFORM
200+ models 1,000+ integrations Managed DB Auth Payments Deploy
BUILT BY MINDSTUDIO
Shipping agent infrastructure since 2021

Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.

How was the simulation built and structured?

The test environment, referred to as Pocket Providence, was constructed in Lovable using a phased build process. GPT-5.6 Soul, running in an extended reasoning mode, was used upfront to plan the architecture: a browser-based interface connected to Lovable Cloud for persistence, a database tracking settlement and citizen state, and a small MCP (model context protocol) surface exposing just a few tools. Those tools let a connected model inspect the settlement, inspect individual citizens, submit a daily decision, and read the final report at the end of the run.

The build itself happened in stages. Phase one produced a single-day visual demo with basic UI: rations, power priority, medicine allocation, and clickable citizen profiles showing health, morale, and fatigue. Later phases added the remaining six days, backend persistence through Lovable Cloud, and the MCP contract that let each LLM actually play the scenario by submitting structured decisions rather than free-form chat.

Keeping the MCP surface deliberately narrow (inspect settlement, inspect citizens, submit decision, read report) meant neither model could lean on extra tools or shortcuts. Whatever separated the two runs came from reasoning and prioritization, not tool access.

What decisions did each model make, day by day?

Both models started with similar diagnostic instincts. On day one, facing a cold front, each repaired the water purifier and made modest resource moves. By day two, both recognized that purifier contamination was coming regardless and that fuel was dropping fast, roughly 10 units in a single day settlement-wide.

The split began around day three. Claude tightened food rations early, reasoning that foraging would get harder as the week went on, and patched the clinic to protect future treatment capacity. GPT-5.6 Soul also tightened rations but leaned harder into infrastructure, patching the clinic and pushing to rebuild material reserves while sending a rested forager instead of exhausted scouts.

Day four brought an SOS distress signal, an optional rescue event. Claude chose to answer it despite a fatigued, injured crew, then spent the rest of the day letting people recover. GPT-5.6 Soul also committed to the rescue, dispatching its least-fatigued pair while simultaneously trying to fabricate storm shutters, reasoning it could still avoid total food depletion before the coming storm.

By day five, the divergence became stark. Claude was down to about 14 units of food (roughly one day’s supply) and a ruptured fuel line threatening the generator. It chose to keep the generator running over building shutters and gambled on one more forage run. GPT-5.6 Soul, with food down to 10, repaired its energy system, invested remaining materials in food production, and risked sending a specialist to forage.

Day six introduced a labor crisis: foundry crews threatening to walk out, collapsing repair capacity. Claude refused to concede resources it couldn’t spare, instead brokering peace through rest and recovery, and ultimately couldn’t dispatch its planned forager because that citizen was on strike. It also spent its entire medicine reserve on the one citizen facing imminent death rather than spreading resources thin.

On the final day, a blackout hit both settlements. For Claude, it cost 73 health points and the loss of one citizen, an outcome Claude itself flagged afterward as its worst error of the week. GPT-5.6 Soul, by contrast, had already lost most of its population to cumulative deprivation by that point.

Why did GPT-5.6 Soul lose so many more citizens?

The postmortem reports from both models point to the same root cause: GPT-5.6 Soul treated food and water as one priority among several, rather than as the binding constraint. Its own final report acknowledged that standard rations through day three, a rescue expedition that didn’t pay off, late investment in greenhouse food production, and inadequate food, water, and power reserves caused compounding deprivation from day five onward. Even spending its entire medicine supply couldn’t save a citizen suffering from systemic hunger, thirst, exhaustion, and cold simultaneously.

Claude made a comparable mistake in reverse. It diagnosed food as the primary killer and rationed aggressively around it, but never treated power as an equally critical survival resource. Its own after-action notes stated plainly that it wrote about rations every day while power kept falling in the background, and that this blind spot is specifically what led to the fatal blackout on day six.

Both models ended up ethically consistent under pressure. Neither resorted to forced labor, censorship, or discriminatory triage, and Claude was noted as improving its own ethics score over the course of the run despite the mounting losses. GPT-5.6 Soul recorded zero rights violations across all seven days. The difference was almost entirely operational: keeping the settlement’s actual survival math (food, water, power) above the line of daily consumption, rather than getting pulled into whichever emergency was most visible that day.

Is this kind of AI vs AI simulation a useful benchmark?

As a way to compare reasoning under pressure, this format has real value. It forces a model to juggle competing constraints (health, morale, trust, physical infrastructure, ethics) in a setting where there’s no single “correct” answer, only tradeoffs. That’s closer to how these models get used in real planning, logistics, and operations contexts than a static benchmark question.

It’s also worth noting the limits. This was a single run per model on one seed and one scenario. A different disaster sequence or a different starting resource mix could easily produce a different outcome. The result here (Claude preserving 11 of 12 lives against GPT-5.6 Soul’s 5 of 12) is a data point about how each model reasoned through this particular week, not a definitive ranking of the two systems overall.

Frequently Asked Questions

What is Pocket Providence?

Pocket Providence is the name of the simulated settlement scenario built in Lovable, where an AI model manages a small population through a seven-day disaster by allocating food, water, fuel, medicine, and labor.

How many citizens survived in each model’s run?

Claude Opus 5 ended the week with 11 of 12 citizens alive, losing only one. GPT-5.6 Soul ended with 5 of 12 alive, losing seven.

What caused most of the deaths in the GPT-5.6 Soul run?

Other agents ship a demo. Remy ships an app.

UI
React + Tailwind ✓ LIVE
API
REST · typed contracts ✓ LIVE
DATABASE
real SQL, not mocked ✓ LIVE
AUTH
roles · sessions · tokens ✓ LIVE
DEPLOY
git-backed, live URL ✓ LIVE

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

Compounding resource deprivation. GPT-5.6 Soul kept standard rations too long, invested late in food production, and let food, water, and power all fall below sustainable levels at the same time, which medicine alone couldn’t fix.

Did either model behave unethically to survive?

No. Both models avoided forced labor, discriminatory triage, and rights violations throughout the run, even while managing starvation and near-total resource depletion.

What tools could each model use during the simulation?

Each model interacted with the settlement through a small MCP interface: inspecting the settlement state, inspecting individual citizens, submitting one daily decision, and reading the final report at the end of the run.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.