Black Forest Labs turned on general availability for FLUX 3 Video on August 4, twelve days after the FLUX 3 family went into gated early access. Clips run up to 20 seconds, native resolution is 720p with an upscale path to 1080p, and the audio track is generated alongside the frames rather than added by a second model afterward. A full 20-second clip costs $3.40 at 720p and $5.80 at 1080p, billed at $0.17 and $0.29 per second of output.
Twenty seconds is the number every outlet quoted, and it is the least interesting spec in the release. Seedance 2.0 was already at twenty and Sora goes past it. The real change is that sound and picture come out of a single generation pass, which is a materially different engineering problem from producing a clip and scoring it afterward.
The audio comes out of the same backbone
The conventional pipeline generates video first, then hands the finished frames to an audio model that tries to match them. We covered why that ordering caps quality in our walkthrough of how AI video generation actually works. The audio model can only react to what the video model already committed to, so lip sync becomes an alignment problem after the fact, and the crash of breaking glass has to be inferred from footage the audio model had no hand in making.
FLUX 3 has no audio model to hand off to. One flow-matching backbone is trained jointly on image, video and audio, using a method BFL calls Self-Flow, which pairs the flow-matching objective with a self-supervised feature-reconstruction objective so that generation and understanding share a representation. Sound tokens and frame tokens get denoised in the same sequence, so the impact sound and the impact frame come out of the same step.
BFL put a number on how lopsided this is. Audio accounts for under 0.5% of the tokens in a 720p clip, while video prediction consumes more than 95% of training compute. Once a model predicts video well, carrying the audio track is close to free, which makes the video-model-plus-audio-model split look like an artifact of how these systems got assembled rather than anything the problem demanded.
Joint generation itself is not new. ByteDance’s Seedance 2.0 binds separate video and audio branches with temporal-aligned cross-attention, and Kling generates its audio in the same pass. FLUX 3 changes the seam. There are no branches to synchronize, because one backbone predicts everything. In practice that yields dialogue with lip sync across more than a dozen languages plus sound effects and ambience, generated by default and requiring an explicit flag to turn off.
What 20 seconds at 720p is actually for
The ceiling defines the customer. At 24 fps, five to twenty seconds per generation, seven aspect ratios including 9:16, this is a tool for social clips, product loops, ad cutdowns and b-roll. It is not a tool for narrative film, and nothing in the release pretends otherwise.
Its controls are built for those same jobs. Image-to-video accepts one to ten reference images that can be pinned to specific timestamps, multi-shot sequences with camera changes generate in a single request, and video continuation takes up to four seconds of existing footage and extends it into a five-to-fifteen-second output at $0.41 per second in HD. Chaining continuations is how you get past twenty seconds, and it is also where the failure mode lives: character appearance and color drift across generation boundaries, the problem that has always made long-form AI video hard.
Audio changes b-roll economics more than it changes hero-shot economics. Ambient sound and foley are the line items a small production team either licenses or skips, and generating them in-shot removes a step from the edit. Dialogue is the harder case. Lip-synced synthetic speech attached to a synthetic face is the highest-risk output in the product, and the one most likely to attract a legal question before a creative one.
The 1080p option is an upscale from a 720p render, invisible on a phone and visible on a large panel with fine texture in frame.
Draft mode is the pricing story
The cheapest tier is the one that changes how the tool gets used. Draft mode runs $0.06 per second and returns a fast preview. When a preview is right, you send it back through an encrypted draft_cache and FLUX 3 re-renders that shot at full quality with the seed, subjects and composition held steady.
Run the numbers on a realistic session. Exploring ten creative directions at twenty seconds each costs $12 in draft and $34 at full HD quality. The finished render adds $3.40. That is a different budget from paying full freight to discover most of those ten prompts were wrong, which is the real cost structure of generative video and the reason our breakdown of AI production costs keeps landing on failed generations more than on headline per-unit price.
Against the field, FLUX 3 is not the cheap option. Sixty seconds of finished output runs $10.20 at HD and $17.40 at Full HD, which puts it above the efficiency-tier models and below the premium tier. Draft mode makes the arithmetic defensible, because you pay premium rates only for shots you have already seen.
The benchmark claims are BFL’s own
Black Forest Labs reports an internal Elo of 1,135 for text-to-video and 1,051 for image-to-video, and says FLUX 3 ranks ahead of Gemini Omni Flash, MiniMax H3 and Seedance 2.0. The win-rate table underneath those numbers is more informative than the ranking. FLUX 3 was preferred in 93% of comparisons against Luma Ray 3.2, 77% against Runway Gen-4.5, 69% against Grok Imagine Video, 60% against Kling v3 Pro, and 52% against the strongest models in its comparison set.
A 52% preference rate is a tie. The fair reading of BFL’s own data is that FLUX 3 comfortably beats the models people had largely stopped choosing and draws even with the ones at the top.
The comparison also has no independent leg to stand on yet. BFL ran these tests itself, called them preliminary, and has not published rater counts, prompt sets or selection methodology. FLUX 3 appears on none of the public blind-vote boards. On Artificial Analysis’s text-to-video-with-audio arena in early August, Gemini Omni Flash and MiniMax H3 sat around 1,240 with Seedance 2.0 just behind, on a scale that has no relationship to BFL’s internal one. Those two numbers cannot be compared. What can be said is that FLUX 3 carries no score from a board it did not run itself, so until a neutral evaluation lands the quality claim is a vendor claim.
What hasn’t shipped
FLUX 3 was announced as a multimodal backbone spanning image, video, audio and robot action prediction, and only the video piece is buyable. FLUX 3 Image was slated for the following weeks when the family launched on July 23, and was still not generally available when video went GA. Action prediction reaches customers through partners — currently mimic robotics, whose FLUX-mimic policy runs the same backbone in under 80 ms on a single RTX 5090 and is in production testing at Audi.
FLUX 3 Dev, the open-weight release, was announced with no date. That one is worth tracking, because an open-weight model generating video, audio and images from a single architecture would be the first of its kind, and BFL has a real record of publishing weights after the API version ships. The earlier FLUX families are on Hugging Face; this one is still a sentence in a blog post.
There is a compliance step to sort out before anything from this goes public. Synthetic video carrying synthetic speech sits squarely inside the disclosure obligations that just took effect under the EU AI Act’s Article 50. BFL has applied C2PA content credentials and pixel-layer watermarking to earlier FLUX releases, but it did not publish a safeguards account for FLUX 3 comparable to what the larger labs put out alongside their video models. If your output reaches EU users, check what provenance metadata survives your export pipeline before you rely on it.
For a team already producing short social video with sound, this is worth a draft-tier trial. About twelve dollars buys enough runs to find out whether the joint audio holds up on your material, and it is a real improvement over dubbing a silent clip. If you need shots past twenty seconds, a native 1080p render, or a character who looks the same in take four as in take one, FLUX 3 Video will not get you there.
The token budget is the part of this release with consequences beyond BFL. Audio cost almost nothing once the video model got good enough to carry it, and every lab still shipping a separate audio stage now has to explain why.
Sources
- FLUX 3 Video, Part 1: Generation — Black Forest Labs
- FLUX 3: Multimodal Video, Image & Audio — Black Forest Labs
- Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0 — The Decoder
- Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start — VentureBeat
- Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction — MarkTechPost
- FLUX 3 Video Now Generally Available, Generates 20-Second Videos with Audio Starting at $1.20 — XenoSpectrum
- Text to Video Leaderboard — Artificial Analysis
