Fal H3 Max Pricing: Free Tier, API Costs and Discount Explained
How much Fal's H3 Max video model costs via API, the current 50% off promo, and how to generate video for free on Fal right now.

What does Fal H3 Max cost right now?
Fal prices H3 Max API generations at 8 cents per second of video at roughly 720p (768 resolution), which works out to about $4.80 per minute of generated footage. Fal is currently running a limited-time promotion that cuts that in half, to 4 cents per second, for the first 14 days. Fal also offers a number of free generations per day through its web interface, no API key or payment required, so anyone curious about the model can try it without spending anything.
TL;DR
- H3 Max costs 8 cents per second via API (about $4.80 per minute) at 768/720p resolution, before any discount.
- A 50% off promo drops the rate to 4 cents per second for the first 14 days, making short experimental clips cheap to test.
- Fal gives users free daily video generations through its site, no credit card needed, with both text-to-video and image-to-video options at 5 or 10 second lengths.
- Running a 24/7 continuous stream at this pricing gets expensive fast, with a rough estimate landing near $4,000 a day for something like an always-on generated channel.
- H3 Max is built on MiniMax’s open-source H3 model but the “Max” version and its Director mode (the auto-regressive continuity layer) are Fal’s own additions, and it’s unclear whether Director will ever be open sourced.
- Speed is the headline feature, with Fal claiming sub-3-second generation for a 5-second 768p clip with audio, which is faster than real time.
One coffee. One working app.
You bring the idea. Remy manages the project.
How does H3 Max pricing actually break down?
The core rate is straightforward: 8 cents per second of output video. A 5-second clip costs 40 cents, a 10-second clip costs 80 cents, and a full minute runs close to $4.80. That’s the standard API price for generating at the 768-ish resolution tier Fal is currently offering.
The promotional rate cuts all of that in half, to 4 cents per second, for a two-week window from launch. During the promo, a 5-second clip drops to 20 cents and a minute of footage costs around $2.40. That’s a meaningful discount if you’re testing workflows, building a demo, or running a batch of short clips to evaluate quality before committing to production use.
Where costs escalate is with continuous or always-on generation. If you wanted to replicate something like an infinite, always-streaming AI channel (a real project some builders have already launched), the math points to something in the neighborhood of $4,000 a day at standard pricing. That’s not a typo, it’s just what happens when you multiply a per-second API cost by 86,400 seconds. For anyone thinking about building a 24/7 generative stream, that number alone makes it clear the current pricing model isn’t built for always-on use cases yet.
Is the free tier actually usable, or just a limited demo?
The free tier on Fal is genuinely usable for casual experimentation. Users get a set number of free video generations per day directly through the Fal website, no API integration required. You can choose between text-to-video or image-to-video, and generate clips at either 5 or 10 seconds. In practice, generations queue up (since everyone hitting the free tier shares capacity), so actual wait times can run longer than the sub-3-second benchmark Fal advertises for the paid API, sometimes closer to 30 seconds depending on load.
Still, for anyone wanting to test what H3 Max produces before touching the API or writing any code, the free tier is the easiest entry point. It’s a reasonable way to sanity-check quality, resolution, and motion behavior on your own prompts before deciding whether the paid tier is worth building around.
What are you actually paying for: speed versus quality?
H3 Max’s main selling point is speed, not necessarily best-in-class visual fidelity. Fal has stated the model can render a 5-second, 768-resolution clip with audio in under 3 seconds, which is faster than the clip’s own runtime. That’s a notable technical jump for a video model, and it’s what enables the real-time and interactive use cases built on top of it (live-prompted streaming channels, choose-your-own-adventure style branching video, and continuous generation loops).
But faster inference generally means trade-offs elsewhere. Side-by-side comparisons between H3 Max and locally-run versions of the open-source H3 model (using acceleration techniques like a “Turbo LoRA”) have shown the local, slower-but-more-deliberate route can produce slightly better results on identical prompts. This tracks with the general “fast, cheap, good, pick two” trade-off common across generative AI tooling. H3 Max isn’t aiming to be the sharpest video model available. It’s aiming to be fast and cheap enough to make real-time, interactive, and streaming use cases possible at all, which is a different goal than maximizing per-clip quality.
What’s the Director feature, and does it cost extra?
Beyond the base model, Fal has a variant called H3 Max Director, which is an auto-regressive mode that keeps up to two minutes of contextual continuity across generations. In practical terms, this means characters, locations, and scene details can stay consistent as new clips generate in sequence, rather than each 5 or 10 second segment starting from scratch. This is what allows the more structured demos (like an interactive Seinfeld-style narrative bit that was floating around online) to maintain some coherence despite being built from short, chained clips.
Director is a Fal-specific layer built on top of the open-source H3 foundation, not something MiniMax released. Standard per-second API pricing still applies to Director-generated clips, though workflows using it may also call small language models to help manage narrative continuity between beats, which can add modest additional cost on top of the video generation itself if you’re building your own interactive layer around it.
Will H3 Max be open sourced?
There’s no official confirmation yet, but there has been discussion suggesting the core H3 Max model could eventually be open sourced, following the pattern MiniMax set with the original H3 model. The Director mode, however, is widely expected to remain proprietary to Fal, since it’s their own addition rather than part of MiniMax’s original release. If the base model does get open sourced, it would likely need to be run locally with your own hardware and inference setup, similar to how the original H3 model is currently used, rather than through Fal’s hosted API pricing.
Frequently Asked Questions
How much does Fal H3 Max cost per second of video?
Standard pricing is 8 cents per second at 768 (roughly 720p) resolution. A current promotion cuts that to 4 cents per second for the first 14 days.
Can I generate video with H3 Max for free?
Yes. Fal offers a limited number of free generations per day through its website, supporting both text-to-video and image-to-video at 5 or 10 second lengths.
How much would a 1-minute video cost?
At standard pricing, roughly $4.80 per minute. During the current 50% off promotion, roughly $2.40 per minute.
Is H3 Max the same as the open-source MiniMax H3 model?
It’s built on that open-source foundation, but H3 Max and its Director continuity feature are Fal’s own hosted, optimized version, distinct from running the base H3 model locally.
Is H3 Max good for building a 24/7 AI video channel?
Not at current pricing. Running continuous generation nonstop works out to an estimated few thousand dollars per day, which makes it impractical for most always-on projects right now.

