Flux 3 Is Here: What Black Forest Labs' New AI Video Model Can Do
Black Forest Labs launches Flux 3, a multimodal AI video generator with 20-second clips and audio. Here's what early access reveals so far.

What is Flux 3?
Flux 3 is the video generation model from Black Forest Labs, released as the follow-up to the image-focused Flux 1 that launched back in August 2024. It’s a multimodal model, meaning it accepts text prompts, image references (up to 10 of them), audio, and video inputs for editing or remixing existing footage. It generates clips up to 20 seconds long, supports standard aspect ratios from 9:16 to 21:9, and produces native audio alongside video. It’s currently in early access via Discord, with resolution capped at 720p at launch and 1080p expected to roll out shortly after.
TL;DR
- Flux 3 launched in early access more than a year after Black Forest Labs first teased a video model alongside Flux 1, making this a long-awaited release for the company.
- Clip length hits 20 seconds, a meaningful jump from the 10-second ceiling that’s become the norm for most competing video models.
- The model is multimodal, accepting up to 10 image references plus audio and video inputs, so it handles editing and remixing tasks, not just fresh generation.
- Early testing suggests reasoning-like behavior, with text-to-video outputs showing solid coherence even on tricky, multi-beat prompts.
- Image-to-video had rough edges in early access, specifically with reference images failing to stick consistently, though this is expected to be ironed out before wider release.
- Open weights are reportedly part of the launch plan, and generation costs are expected to land below Seedance pricing, though no official pricing has been announced.
- It’s positioned as a complement to existing tools like Kling and Seedance, not an outright replacement, at least based on day-one testing.
Plans first. Then code.
Remy writes the spec, manages the build, and ships the app.
How does Flux 3 compare to Seedance and Kling?
The honest answer from early hands-on testing: it’s close, but not a clean win either way. Flux 3 matches or beats the field on raw clip length (20 seconds versus the 10-second norm) and on multimodal flexibility, since it accepts far more reference inputs than most rivals. Prompt coherence in text-to-video tests looked strong, holding up even on deliberately difficult prompts designed to stress-test a model’s ability to track multiple actions and scene changes.
Where it currently falls short is kinetic energy. Side-by-side testing against a demanding Seedance prompt (a rapid-cut, hyperpop-style sequence) showed Flux 3 hitting every element requested in the prompt, but without the same frantic camera work, fisheye trick shots, or fast action rhythm that Seedance produced. For high-energy action sequences, Seedance currently has an edge. For dialogue-driven or dramatic scenes, Flux 3’s outputs held up well, including one community-generated clip with convincing rapid-fire dialogue and believable delivery.
The takeaway from testing so far: Flux 3 isn’t a replacement for Seedance or Kling in most creator workflows, but it’s a strong enough addition to sit alongside them, particularly for creators who need longer clips or want to feed in audio and multiple image references at once.
What are the early access limitations?
The most notable issue during early testing was image-to-video generation, specifically getting image references to actually attach and influence the output consistently. This happened through the Discord-based early access interface, which introduces its own friction (finding and crediting community outputs in a busy Discord server is genuinely difficult). It’s a launch-day issue rather than necessarily a model limitation, and it’s reasonable to expect it gets resolved as the model moves toward broader release.
Resolution is also capped at 720p for now, with 1080p expected soon. And since this is day one of the model being available even in limited form, there hasn’t been time to fully map out its edge cases, strengths, or failure modes. Every early read on a brand-new model comes with the caveat that updates are coming fast.
Does Flux 3 do audio generation well?
Yes, audio generation is a built-in feature rather than a bolt-on. A text-to-video test using a simple prompt (asking for a trailer for a fictional Netflix show called “Flux 3: Return to the Black Forest”) produced a coherent trailer-style clip with audio, upscaled afterward in Topaz to compensate for the 720p output resolution. The result wasn’t flawless, but it held together well enough to suggest the model is doing some form of reasoning or planning across the full clip rather than generating frame-by-frame without a broader sense of structure.
A separate image-to-video test, a recurring benchmark involving an FBI agent drinking coffee in a Pacific Northwest diner, showed the extra 5 seconds of runway (20 seconds versus the usual 15) giving background details and secondary dialogue more room to land, including a bit part character who has historically been cut off mid-sentence in shorter clips from other models.
Is Flux 3 open source?
No official confirmation yet on open weights, but the direction points that way. According to reporting picked up alongside the early access period, open weights appear to be part of the launch plan, meaning a version of Flux 3 could eventually be available for local or self-hosted use, similar to how earlier Flux image models were released. No pricing has been announced for the hosted version either, but expectations are that generation costs will land below Seedance’s pricing, which is on the higher end of the current video model market.
Is Flux 3 worth trying right now?
For creators already deep in a Kling or Seedance workflow, Flux 3 is worth testing as a complement rather than a wholesale switch. The 20-second clip length alone is useful for dialogue scenes or slower dramatic beats that get cut short at 10 seconds elsewhere. The multimodal input support (image, audio, video) also opens up editing and remixing use cases that pure text-to-video or image-to-video models don’t handle as directly.
The rough edges are real, though. Image-to-video reference handling wasn’t reliable in early access, resolution is currently limited, and fast-action sequences don’t have the same visual energy as top competitors. Anyone evaluating it today should treat it as a preview of where the model is headed rather than a finished product, and expect meaningful changes as it moves from early access Discord testing to a public release with confirmed pricing and higher resolution output.
Frequently Asked Questions
What is Black Forest Labs?
Black Forest Labs is the company behind the Flux family of AI image generation models, with Flux 1 launching in August 2024. Flux 3 marks their entry into video generation, a capability that was teased at the original Flux 1 launch.
How long are Flux 3’s video clips?
Flux 3 can generate clips up to 20 seconds long, longer than the roughly 10-second ceiling common across most current AI video models.
Does Flux 3 support image-to-video generation?
Yes, though early access testing showed inconsistent results with image references failing to reliably attach to prompts. This is considered a launch-stage issue likely to be fixed before wider release.
Will Flux 3 be open source?
Open weights reportedly are part of the launch plan, though no confirmed release date or pricing has been announced as of early access.
How does Flux 3 compare to Seedance for action scenes?
Seedance currently has an edge on fast, kinetic action sequences with rapid camera movement, while Flux 3 handled dialogue-heavy and dramatic scenes well in early testing, hitting prompt requirements without the same visual intensity on action-focused prompts.
