Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace

How to Build a Multiplayer AI Game With Claude Opus 5.5

A practical breakdown of building an Attack on Titan x Minecraft multiplayer game with Claude Opus 5.5, voice prompting, and live playtesting.

Edited by Luis Chavez-Mattos, Director of Product RSS
How to Build a Multiplayer AI Game With Claude Opus 5.5

How do you actually build a multiplayer game with an AI coding agent?

You start with a working single-player prototype, pick a capable model to run the build, then iterate through voice-prompted rounds of playtesting and revision until the mechanics feel right. That’s the workflow one creator used to turn a single-player “Attack on Titan meets Minecraft” prototype into a playable multiplayer game using Claude Opus 5.5 inside an AI coding agent. The process wasn’t a one-shot prompt. It was a loop: build, play, narrate what’s wrong out loud, repeat.

TL;DR

  • Model choice changes both quality and cost, with Haiku 5.5 producing a workable but rougher prototype for around $5, while Sonnet 5.5 and Opus 5.5 landed in a similar $24 to $25 range for a noticeably more polished result.
  • Voice prompting beats typing for ideation speed, letting the builder dictate around 200 words per minute of design direction instead of typing at a fraction of that speed.
  • The build started from an existing single-player prototype, which was then explicitly reframed as a multiplayer game spanning two playable factions: Survey Corps members and Titans.
  • Playtesting drove the design more than upfront planning did, with mechanics like fall damage, ODM gear braking, and nape-hit indicators discovered by actually playing the game rather than from a spec document.
  • Game balance was handled through explicit constraints, such as capping the Colossal Titan’s power so it couldn’t instantly end a match.
  • The agent was told to ask for clarification rather than guess, which kept the AI from making large unilateral design decisions during a long, multi-part build.
  • Multiplayer-specific systems were added deliberately, including kill/death boards, bot fallback when teams are short on players, and League of Legends style escalating respawn penalties.

Other agents start typing. Remy starts asking.

YOU SAID "Build me a sales CRM."
01 DESIGN Should it feel like Linear, or Salesforce?
02 UX How do reps move deals — drag, or dropdown?
03 ARCH Single team, or multi-org with permissions?

Scoping, trade-offs, edge cases — the real work. Before a line of code.

What tools were used to build the game?

The builder ran Claude Code with Opus 5.5 inside a terminal environment, using an app called Orca to manage multiple AI agent threads at once. Alongside the Opus 5.5 thread, a separate thread ran Codex with a different model, which let the builder compare approaches or parallelize work across agents. None of this requires a specialized setup: the same workflow is possible with the standard Claude desktop app or the Codex app, since the point isn’t the interface but the ability to hold a long, iterative conversation with the model while it writes and rewrites code.

The other key tool was a voice transcription layer. Instead of typing instructions, the builder spoke them directly into the agent, treating the first prompt almost like a design monologue: faction choices, movement mechanics, combat feel, respawn pacing, and balance concerns, all delivered in one long take and then dumped into the agent as a single detailed brief.

Why does model choice matter so much for AI-built games?

Different model tiers produced meaningfully different starting points. In the creator’s informal benchmark, Haiku 5.5 built a working prototype, complete with shadows and animations and a basic fall-damage system, for roughly $5. Sonnet 5.5 and Opus 5.5 cost notably more, landing around $24 to $25, but produced a more refined initial game. That gap matters because the first version acts as the seed for everything after it: a shakier prototype means more rounds of fixes before the core gameplay even feels right, while a stronger prototype lets later iteration focus on multiplayer-specific features instead of basic jank.

The practical implication is that model selection is a cost-versus-quality decision made upfront, not something to revisit mid-build. Once Opus 5.5 was chosen as the production model, the entire multiplayer conversion happened in that same thread.

How does voice prompting change the workflow?

Typing forces a builder to compress their intent into short, efficient instructions. Talking doesn’t. The creator’s pitch for voice prompting is simple: most people type somewhere between 40 and 80 words a minute, but can speak around 200. For a creative, exploratory task like describing game feel, that speed difference matters, because it lets the builder dump a large, loosely structured wishlist (faction mechanics, fuel drops, nape-hit indicators, respawn pacing, character selection) in one pass without the friction of typing slowing down the thought process.

This also changed the shape of the prompts themselves. Rather than a tightly scoped feature request, the initial voice prompt read more like a design document spoken aloud, covering combat balance, traversal mechanics, competitive incentives, and even specific character references from the source material, all in a single continuous instruction.

What design decisions shaped the gameplay?

A few choices stand out as deliberate balance calls rather than default AI output:

Remy doesn't build the plumbing. It inherits it.

Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.

200+
AI MODELS
GPT · Claude · Gemini · Llama
✓
1,000+
INTEGRATIONS
Slack · Stripe · Notion · HubSpot
✓
MANAGED DB
AUTH
PAYMENTS
CRONS

Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.

The Colossal Titan was explicitly constrained so it couldn’t become a one-click win condition, since a Titan with overwhelming power would undermine competitive multiplayer matches. Respawning was tuned to feel like a real setback rather than a casual reset, inspired by how games like League of Legends lengthen respawn timers as a match progresses. Fuel for the ODM (omnidirectional mobility) gear was made a limited resource with random drops scattered across the map, so players couldn’t swing indefinitely without managing resources. And when teams were short on human players for the Titan side, the design called for bots and NPCs to fill the gap, keeping matches playable even without a full lobby.

These are the kinds of decisions that separate a tech demo from a game people will actually keep playing: constraints that create tension and consequence rather than just adding features.

Why did playtesting matter more than planning?

The most notable part of the workflow is what didn’t happen: there was no exhaustive design document written before building began. Instead, the builder played the game after each major round of changes and used what felt wrong as the next set of instructions.

That’s how several core mechanics got discovered. Swinging into a surface using the ODM gear had no downside, which felt unrealistic for a game about fragile human soldiers up against giants. Falling from height had no consequence either. Playing exposed both problems immediately, in a way a planning document likely wouldn’t have anticipated. The fix became an explicit instruction back to the agent: players should take damage from high falls and from slamming into obstacles, and only a well-timed braking maneuver with the ODM gear should let a player land safely from a big drop.

That braking mechanic itself went through iteration, too. An early version required players to start braking too early to avoid damage, which didn’t feel right, leading to a follow-up conversation about tuning it, possibly with a dedicated key for grounding quickly.

The underlying principle: when an AI model can generate most of the implementation cheaply, the bottleneck isn’t planning, it’s noticing what feels off and describing it precisely. Playing the game is a faster way to generate that feedback than trying to anticipate every issue on paper.

Is this workflow realistic for other builders?

It’s realistic for anyone comfortable treating an AI coding agent as a true collaborator rather than a one-shot code generator. The workflow depends on being willing to go back and forth repeatedly: build a chunk, play it, describe what’s wrong, let the agent revise, play again. It also depends on giving the agent explicit permission to ask questions instead of guessing when instructions are ambiguous, which keeps a long, multi-feature build from drifting away from the intended design.

The costs are real but not exotic. Cloud-based model usage for a build like this runs from single digits to around the mid-twenties in dollars just for the initial prototype, before accounting for the additional iteration needed to reach multiplayer. That’s a meaningful but accessible cost for a hobbyist project, and far cheaper than traditional game development.

Frequently Asked Questions

What models were compared for building the game?

Haiku 5.5, Sonnet 5.5, and Opus 5.5 were each given the same task to generate an initial version of the game, which let the builder compare output quality against cost before settling on Opus 5.5 for the full multiplayer build.

How much does it cost to build a game like this with AI?

The initial single-player prototype cost around $5 with Haiku 5.5, and around $24 to $25 with either Sonnet 5.5 or Opus 5.5. Those figures cover just the first build, not the full iteration needed to add multiplayer features.

What is voice prompting and why use it for coding agents?

Voice prompting means dictating instructions to an AI agent through a transcription tool instead of typing them. It’s faster for getting detailed, exploratory ideas out, since speaking averages around 200 words a minute compared to roughly 40 to 80 words a minute of typing.

Did the builder write a full design document before starting?

No. The approach favored playing the game after each round of changes and describing problems as they came up, rather than trying to anticipate every mechanic in advance through planning.

What tools were used to run multiple AI agents at once?

The builder used an app called Orca to manage separate threads, running Claude Code with Opus 5.5 in one tab and Codex with another model in a different tab, though a single standard app like Claude desktop or Codex works too.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.