Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
run Iris 3B locallyIris 3B VRAMlocal image generation

How to Run Iris 3B Locally: VRAM Needs and Setup Explained

A practical guide to running Iris 3B locally, including its ~28GB VRAM footprint, Gradio demo setup, and real-world generation times.

Edited by Luis Chavez-Mattos, Director of Product RSS
How to Run Iris 3B Locally: VRAM Needs and Setup Explained

What is Iris 3B and why does it run differently from other image models?

Iris 3B is a 3 billion parameter image generation model from Spirit Labs that generates images by painting pixels directly, with no variational autoencoder and no latent space involved. Most popular image generators, including Flux and Qwen Image, compress an image into a latent representation first and decode it back into pixels at the end. Iris skips that step entirely. It also doubles as a general vision model: the same weights, fine-tuned, handle depth estimation and image upscaling, both of which are included in the demo setup that creator Fahad Mirza walked through on his channel.

TL;DR

  • Iris 3B needs roughly 28GB of VRAM to run locally, putting it in the same bracket as Flux and Qwen Image despite having far fewer parameters than some of those systems.
  • Installation is handled through Spirit Labs’ model card on Hugging Face, with a Gradio demo layered on top for a browser-based interface.
  • Generation is slow by modern standards, taking roughly one to three minutes per image because the model works in pixel space instead of a compressed latent space.
  • The same checkpoint, fine-tuned, supports depth estimation and image upscaling in addition to text-to-image generation, all inside the same demo.
  • Image quality is strong on texture and fine detail (skin pores, beard hairs, fabric weave, rug patterns) but inconsistent on cultural and anatomical accuracy, sometimes misreading what a prompt is actually describing.
  • The architecture relies on a frozen Qwen3-VL 4B text encoder, dual-stream and single-stream transformer blocks, and a small decoder stack that turns tokens back into raw pixels without any VAE.

One coffee. One working app.

You bring the idea. Remy manages the project.

WHILE YOU WERE AWAY
✓Designed the data model
✓Picked an auth scheme — sessions + RBAC
✓Wired up Stripe checkout
✓Deployed to production
Live at yourapp.msagent.ai

How much VRAM does Iris 3B actually need?

Based on hands-on testing, Iris 3B consumes just under 28GB of VRAM when running locally. That number is notable because Iris is a 3 billion parameter model, which is small by current image-generation standards, yet its VRAM footprint lands in the same range as considerably different architectures like Flux and Qwen Image. The explanation is the pixel-space approach: without a VAE compressing the image into a smaller latent grid, the model has to hold and process full-resolution pixel representations throughout generation, which keeps memory usage high regardless of parameter count.

Practically, this means Iris 3B is not a model you can expect to run on consumer cards with 8GB or 12GB of VRAM. You’re looking at hardware in the 24GB-plus class, think RTX 3090, RTX 4090, or a comparable data center card. If you don’t have that locally, renting a GPU by the hour from a cloud provider is a reasonable workaround for testing the model before committing to local hardware.

How do you install and run Iris 3B locally?

Setup follows the standard pattern for Hugging Face-hosted models: Spirit Labs provides installation instructions directly on the model card, and a Gradio interface sits on top to give you a browser-based UI instead of requiring you to script every generation call. The general flow looks like this:

  1. Pull the model and dependencies as specified on the Iris 3B model card.
  2. Launch the Gradio demo, which exposes text-to-image generation, depth estimation, and upscaling as separate tabs or modes.
  3. Set your prompt, along with the number of inference steps and guidance scale. The creator kept these at the values Spirit Labs recommends in the model card for consistency when comparing outputs to the official examples.
  4. Submit and wait. Unlike fast latent-diffusion models that can return an image in seconds, Iris 3B takes noticeably longer per generation.

Because there’s no separate VAE decoding step, the entire generation process happens in pixel space from start to finish, which is part of why the setup is simple (fewer moving parts, no separate autoencoder to manage) but also why it’s slower.

How long does Iris 3B take to generate an image?

Generation time runs roughly one to three minutes per image on the hardware used in testing. That’s considerably slower than typical latent diffusion models, which often return results in seconds to low tens of seconds on similar hardware. The tradeoff is architectural: operating directly in pixel space avoids the information loss that comes from compressing into a latent grid and decoding back out, but it also means the model is doing more computational work per pixel rather than per latent unit.

For anyone evaluating Iris 3B for a workflow, this speed difference matters more than the VRAM number in some ways. A 28GB VRAM requirement is a one-time hardware decision. A one-to-three-minute generation time is a recurring cost every time you want an image, which changes how practical the model is for iterative prompt refinement or batch generation.

Is Iris 3B worth running locally?

Remy doesn't write the code. It manages the agents who do.

R
Remy
Product Manager Agent
Leading
Design
Engineer
QA
Deploy

Remy runs the project. The specialists do the work. You work with the PM, not the implementers.

It depends on what you need. The quality on fine detail is genuinely strong. Testing across a wide range of prompts, skin texture, individual beard hairs, woven fabric, rug patterns, and frost on eyelashes all rendered with sharp, convincing detail. Depth estimation output was described as accurate down to finger angles and bone structure, and the upscaling mode produced visibly sharper results on low-resolution source images, with fine details like chipped paint becoming legible after processing.

But accuracy on subject matter and cultural specifics was inconsistent. A prompt for an elderly Indigenous Australian elder produced a result that didn’t actually look Australian Indigenous, and a kangaroo-skin cloak came out looking more like spotted big cat fur. A prompt combining space and cowboy imagery in one frame almost entirely dropped the space element. A Palestinian konafa dessert prompt produced something closer to baklava. Other prompts, covering Siberian Yakutia, Baloch elders in Kalat, Indonesian rendang with traditional Rumah Gadang architecture, Chilean cueca dancing, Zulu dance movement, Munich beer festival scenes, and Persian rug weaving, came out convincingly, with good texture, lighting, and cultural detail.

The pattern suggests Iris 3B is better at rendering materials, textures, and lighting than at correctly interpreting less common cultural or anatomical specifics. If your use case leans on photorealistic texture work, depth mapping, or upscaling, it performs well. If you need reliable cultural or factual accuracy across diverse global subject matter, expect to need multiple generations and some prompt iteration.

How does the architecture affect local performance?

Iris 3B encodes a text prompt using a frozen Qwen3-VL 4B encoder paired with a small text adapter, then splits the noisy image into 16x16 pixel patches. The main network has eight dual-stream blocks, where text and image tokens carry separate weights, followed by 16 single-stream blocks that process both together with time-step conditioning applied at every block. Each block uses grouped query attention with a sigmoid gate. At the end, four small blocks take semantic tokens from the main trunk and convert each patch back into raw pixel values.

This structure explains both the VRAM usage and the generation speed. There’s no VAE to compress the problem down to a smaller latent space, so the full pixel-patch representation has to move through the entire dual-stream and single-stream stack. More blocks processing more pixels directly translates into more VRAM held during generation and more wall-clock time per image, which is the tradeoff anyone running this locally needs to plan around.

Frequently Asked Questions

How much VRAM does Iris 3B require to run locally?

Just under 28GB of VRAM, putting it in the same range as considerably larger pixel-to-image systems like Flux and Qwen Image, despite Iris having only 3 billion parameters.

How long does it take to generate one image with Iris 3B?

Roughly one to three minutes per image, which is slower than typical latent diffusion models due to the lack of a compressed latent space.

Can Iris 3B do anything besides text-to-image generation?

Yes. The same fine-tuned model handles depth estimation and image upscaling, both accessible in the same Gradio demo used for text-to-image generation.

Does Iris 3B use a VAE or latent space like Stable Diffusion or Flux?

No. Iris 3B generates images by working directly in pixel space, cutting images into 16x16 patches and reconstructing pixels at the end without any variational autoencoder or latent representation.

Is Iris 3B reliable for culturally specific or detailed prompts?

Cursor
ChatGPT
Figma
Linear
GitHub
Vercel
Supabase
goremy.ai

Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

Results are mixed. It renders fine textures, lighting, and materials well, but can misinterpret specific cultural, anatomical, or regional details, as seen in tests involving Indigenous Australian, Palestinian, and mixed-theme American prompts.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.