Iris 3B: The Image Generator That Skips Latent Space Entirely
Iris 3B generates images directly in pixel space with no VAE, and the same weights handle depth estimation and upscaling too.

What is Iris 3B?
Iris 3B is a 3 billion parameter image generation model from Spirit Labs that paints images directly in pixel space, skipping the variational autoencoder (VAE) and latent space step that nearly every modern image generator, including Flux and Qwen Image, relies on. The same model, fine-tuned, also works as a general vision learner capable of depth estimation and image upscaling, which makes it less of a single-purpose generator and more of a shared vision backbone with multiple jobs.
TL;DR
- Pixel-space generation means Iris 3B predicts raw pixels directly instead of denoising a compressed latent representation, which is the approach almost every mainstream diffusion image model uses.
- No VAE anywhere in the pipeline sets it apart architecturally from Flux and Qwen Image, even though all three models land in a similar VRAM footprint when run locally.
- One checkpoint, three jobs: the same underlying model, fine-tuned per task, does text-to-image generation, depth estimation, and image upscaling.
- Text conditioning comes from a frozen Qwen3-VL encoder (around 4 billion parameters) paired with a small adapter, rather than training a text encoder from scratch.
- Image quality holds up well on fine detail: skin texture, pores, individual hairs, bead work, fabric folds, and small repeating patterns like woven rugs come through with real sharpness.
- It struggles with ethnic and cultural specificity in faces, sometimes nailing the costume and setting of a prompt while missing the actual look of the people being described.
- Generation is slow, taking roughly one to three minutes per image, which is noticeably longer than typical latent diffusion models at similar VRAM usage.
Remy doesn't build the plumbing. It inherits it.
Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.
Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.
How does pixel-space generation actually work?
Most image generators built in the last few years, including Flux and Qwen Image, use a VAE to compress an image into a smaller latent space before running the diffusion or flow-matching process. The model denoises in that compressed space, and a decoder expands the result back into full-resolution pixels at the end. This compression step is what makes those models fast enough to run on consumer GPUs: working in latent space means processing far fewer numbers than working on raw pixels.
Iris 3B drops that step. According to its architecture, the model takes a text prompt, encodes it with a frozen Qwen3-VL 4B encoder plus a small text adapter, and then cuts the noisy image directly into 16x16 pixel patches. There is no intermediate latent representation at any point. The main network trunk has eight “dual stream” blocks, where text and image tokens carry separate weights, followed by sixteen “single stream” blocks that process text and image together, with time step conditioning injected at every block. Each block uses grouped query attention with a sigmoid gate. At the end, four smaller blocks take the semantic tokens produced by the trunk and convert each patch back into raw pixel values.
The practical effect of removing the VAE is that the model never loses information to compression artifacts from an autoencoder, and the same pixel-level representation it learns for generation transfers naturally to tasks like depth estimation and upscaling, since those tasks are also fundamentally about predicting pixel values rather than manipulating a latent code.
Why skip the VAE if it works so well for everyone else?
VAEs are efficient, but they’re also a layer of lossy compression. Everything the diffusion model generates has to pass through a decoder trained separately, and whatever that decoder can’t reconstruct faithfully becomes a ceiling on image quality, especially for fine, repeating detail like woven patterns, individual hair strands, or small text.
Iris 3B’s pixel-space approach appears to pay off specifically in that kind of fine detail. In testing shown in a walkthrough of the model, images of a Persian rug weaver in Isfahan showed tight, repeating pattern work, and portraits consistently rendered sharp skin texture, visible pores, individual beard hairs, and the weave of woven fabric. A test image of sky lanterns rising over a festival in Thailand showed visible paper folds on the lanterns themselves, a level of small-scale detail that’s often where latent-space models soften out.
The tradeoff is speed. Working directly on pixels instead of a compressed latent means more computation per image, and generation reportedly takes one to three minutes per image, noticeably slower than typical latent diffusion models.
Is Iris 3B worth running locally?
For people experimenting with local image generation, Iris 3B is interesting mainly because of what it demonstrates architecturally, not because it’s the fastest or most accurate option available. VRAM consumption lands under 28 gigabytes, which is comparable to Flux and Qwen Image despite the very different internal architecture. That’s a notable data point: dropping the VAE didn’t meaningfully increase memory requirements, even though it clearly increased generation time.
Where it’s genuinely compelling is the multi-task angle. The same 3 billion parameter model, with different fine-tuning, handles depth estimation and image upscaling in addition to text-to-image generation. In testing, the depth estimation output correctly separated fine structures like finger joints and hand angles, and the upscaling pass measurably sharpened low-resolution, blurry source images, making chipped paint and fine texture visibly crisper.
Where it falls short is cultural and ethnic specificity in people. A prompt for an elderly Indigenous Australian elder produced a face that didn’t read as Indigenous Australian at all, even though it correctly rendered requested details like red ochre paint, a woven headband, and a white beard. A cloak meant to evoke kangaroo skin came out looking more like spotted big cat fur. Similar gaps showed up elsewhere: a prompt combining a NASA astronaut and a Texas cowboy in one frame mostly dropped the space theme entirely, and a Palestinian knafeh dessert prompt produced something closer to baklava than actual knafeh. Crowd scenes, like devotees at a Chhath Puja festival in Patna, rendered convincing haze, architecture, and water reflections, but individual faces in the crowd tended to look repetitive rather than distinct.
What does this mean for people building with image models?
Iris 3B is a useful reminder that the VAE-plus-latent-space formula, while dominant, isn’t the only way to build a working image generator. A pixel-space model built around a frozen vision-language encoder (Qwen3-VL 4B here) and a dual-stream/single-stream transformer trunk can produce genuinely sharp, detailed output and double as a depth estimator and upscaler using the same weights, just fine-tuned differently.
For builders, the practical takeaway is less about swapping in Iris 3B for production work and more about what the architecture signals: pixel-space generation is viable at the 3 billion parameter scale, VRAM cost isn’t necessarily higher than latent-space competitors, and a single well-trained vision backbone can plausibly serve multiple downstream vision tasks instead of needing separate specialized models for generation, depth, and upscaling.
Frequently Asked Questions
What makes Iris 3B different from Flux or Qwen Image?
Iris 3B generates images directly in pixel space with no variational autoencoder and no latent space at any point in the pipeline. Flux and Qwen Image both compress images into a latent representation first and denoise there before decoding back to pixels.
How much VRAM does Iris 3B need to run?
Local testing showed VRAM consumption under 28 gigabytes, which is in the same range as Flux and Qwen Image despite the different architecture.
Can Iris 3B do more than generate images?
Yes. The same model, fine-tuned for different tasks, also performs depth estimation and image upscaling, and both were demonstrated working directly in the model’s local demo.
Why is Iris 3B slower than other image generators?
Working directly on raw pixels instead of a compressed latent space requires more computation per image. Generation reportedly takes roughly one to three minutes per image.
Does Iris 3B handle cultural and ethnic detail accurately?
Results were mixed. It handled fine textures, traditional dress, and architectural detail well in many cases, but struggled with accurately representing specific ethnic facial features, such as an Indigenous Australian elder prompt that produced a face and clothing texture that didn’t match the intended subject.