FLUX 3 Video went generally available on August 4, 2026, twelve days after Black Forest Labs first announced the FLUX 3 family. The release ships clips up to 20 seconds, audio generated alongside the frames rather than dubbed on afterwards, and a published price list with four full-render rates plus two draft rates — not the single flat number that will end up in most launch coverage.
The distinction matters more than it sounds. FLUX 3 was announced in late July as a multimodal model spanning image, video, audio and action; what actually opened on August 4 is the video generation leg, via the BFL API and select partners. FLUX 3 Image and an open-weight FLUX 3 Dev variant remain announced roadmap with no dates attached. If you are budgeting or building against this release, you are budgeting against one modality, not the family.
This guide covers what the GA actually includes, the mode-specific limits the launch post glosses over, the full six-tier price structure with our own per-minute conversion so it can be compared against the rest of the field, and an honest read on what the quality claims do and do not establish. Every price is labelled by the surface it came from, because two of them disagree.
- 01GA covers the video leg, not the whole FLUX 3 family.Video generation opened on August 4, 2026 via the BFL API and select partners. FLUX 3 Image and the open-weight FLUX 3 Dev variant remain announced roadmap items with no published dates.
- 02There are six published prices, not one.BFL lists $0.17/s and $0.29/s for text-to-video and image-to-video at HD and Full HD, $0.43/s and $0.54/s for video continuation, plus Draft Mode at $0.06/s and $0.12/s. The spread from cheapest to dearest is 9×.
- 03The 20-second ceiling is mode-specific.Text-to-video and image-to-video run 5 to 20 seconds. Video continuation caps at 5 to 15 seconds and takes up to 4 seconds of existing video and audio as its input. Do not apply the headline number uniformly.
- 04The quality claims are vendor-stated, and narrower than they look.BFL reports an internal human-rater preference win on text-to-video and an explicit tie with Seedance 2.0 on image-to-video. No independent video Arena had scored FLUX 3 Video at the time of writing.
- 05Draft Mode is the real cost lever.A draft costs about a third of an HD full render and about a fifth of a Full HD one, and draft-enhance re-renders the approved shot from the draft cache bundle rather than reinterpreting the prompt.
01 — What Went GAGeneral availability for the video leg, not the family.
Black Forest Labs made FLUX 3 Video generation generally available starting August 4, 2026, through the BFL API and select partners. BFL also states the model went live on Krea the same day; treat that as a vendor statement — no dedicated Krea announcement was locatable at the time of writing, so it is not independently corroborated here.
The headline capabilities are straightforward. Clips run up to 20 seconds. Output is released at HD, meaning 720p natively, with Full HD available through upscaling. Audio is generated alongside the video rather than added in a second pass. And three request modes — text-to-video, image-to-video, and video continuation — share a single flux-3-video endpoint with the same request shape.
On OpenRouter the model is listed as black-forest-labs/flux-3-video, created August 4, 2026, with Black Forest Labs as the only provider — no routing or fallback providers at the time of writing. That single-provider status is worth noting if you route production traffic through an aggregator for redundancy: for this model, right now, there is no second provider to fail over to.
This is a different event from the FLUX 3 announcement in late July, which introduced the model family and framed video as gated early access without pricing. If you want that context, our coverage of the July FLUX 3 announcement covers it; this piece deliberately does not restate it.
Text-to-video
Prompt in, clip out. Supports multi-shot generations — multiple scenes and camera angles inside a single request, kept coherent as one sequence rather than stitched together afterwards.
Image-to-video
One image can act as a start frame, or up to ten can storyboard the clip. Passing a [seconds, image] pair pins a given keyframe to an exact timestamp inside the generation.
Video continuation
Feed up to four seconds of existing video and audio, describe what happens next, and the model carries camera movement, dialogue and audio across the seam. Priced separately and higher.
auto, resolution defaults to hd, and audio generation defaults to on. If you do not want a soundtrack, you have to say so.02 — Modes & LimitsOne endpoint, three modes, different ceilings.
The single-endpoint design is genuinely convenient — the same request shape covers all three modes, so switching from a text-driven clip to an image-driven one is a payload change, not an integration change. But the specifications differ per mode in ways the launch post does not spell out, and getting them wrong shows up as failed requests or a blown budget rather than a warning.
The 20-second figure applies to text-to-video and image-to-video. Video continuation is capped at 5 to 15 seconds of output, from up to 4 seconds of supplied footage. Output covers seven aspect ratios — 21:9, 2:1, 16:9, 4:3, 1:1, 3:4 and 9:16 — all at 24 frames per second.
Text-to-video & image-to-video
Both generation modes run a 5 to 20 second range. Video continuation is a different specification entirely: 5 to 15 seconds of output, generated from at most 4 seconds of supplied video and audio.
Image-to-video references
A single image acts as a start frame; up to ten storyboard the clip. A [seconds, image] pair pins a keyframe to an exact timestamp, which is what makes mid-clip control possible rather than just first-and-last-frame steering.
Aspect ratios at 24 fps
21:9, 2:1, 16:9, 4:3, 1:1, 3:4 and 9:16, all rendered at 24 frames per second. Aspect ratio defaults to auto, meaning the model fits the framing to the content unless you pin it.
03 — Native AudioSound generated with the frames, not after them.
Native audio is the capability BFL leads with, and the framing is specific: sound effects, ambience and dialogue are produced together with the frames, not generated separately and aligned afterwards. The parameter that controls it defaults to on.
Dialogue is lip-synced and multilingual. The launch post names roughly thirteen languages and dialect groups explicitly — English in various dialects, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi and Punjabi — followed by “and more.” Read that as about thirteen named with an open tail, not as a certified count; BFL has not published a definitive language list.
“The model generates clips up to 20 seconds long in HD resolution, with Full HD output via upscaling and native audio created alongside the video.”— Black Forest Labs, FLUX 3 Video launch post, August 4, 2026
The second capability worth separating out is multi-shot generation. FLUX 3 Video can produce multiple scenes and camera angles inside a single generation while keeping the sequence coherent. That is a single-request capability, not an editing workflow — which is a meaningful difference for anyone who has built a pipeline that generates shots individually and then fights continuity drift in the edit.
Here is the trade-off nobody puts in a launch post. Joint audio and video generation removes an entire production step, and it also removes an entire control surface. When the soundtrack is baked in at generation time, you cannot swap the voice, re-time the music, or replace a sound effect without regenerating the clip — and regenerating means paying again and accepting new frames. For social-first content and explainer video that is a clean win. For work where audio is licensed, legally reviewed, or brand-locked, the practical pattern is to generate with audio off and keep the sound design in a pipeline you control.
BFL also positions the model for documentary and short-form educational content generated from short prompts, citing pretraining world knowledge combined with real-time grounding. That is vendor positioning rather than a benchmarked capability, and it is the claim most worth testing on your own prompts before you build a content line on it.
04 — PricingSix published prices, and the one everyone will quote.
The number that will travel is $0.17 per second. It is real, and it is the cheapest full-quality tier: text-to-video or image-to-video at HD. It is also one of six published rates. Full HD full renders go to $0.29 per second. Video continuation is priced on its own ladder at $0.43 per second HD and $0.54 per second Full HD on BFL list. Draft Mode sits underneath everything at $0.06 per second for text-to-video and image-to-video, and $0.12 per second for continuation.
Run the spread and the shape of the pricing becomes obvious. From the cheapest published rate to the dearest is a factor of nine. Even restricting to full renders, continuation at Full HD costs 3.2× what an HD text-to-video second costs. A pipeline that treats “FLUX 3 Video” as one line item in a spreadsheet will misprice itself by multiples depending on which modes it actually leans on.
Draft Mode deserves its own paragraph because it is the lever that changes unit economics rather than shaving them. A draft returns a fast, low-cost preview of the prompt; once a draft is approved, draft-enhance renders the same subjects, composition and motion at full quality from the draft cache bundle rather than reinterpreting the prompt from scratch. BFL describes drafts as costing about a third of a full render, which is almost exactly right against the HD tier and understates the saving at Full HD.
$0.06 against $0.17 per second
Our arithmetic on BFL list prices. This is the tier where the vendor’s “about a third” description is exactly right — a draft second costs 35% of an HD full-render second.
$0.06 against $0.29 per second
Against the Full HD tier a draft is closer to a fifth than a third, because the draft rate is flat while the full-render rate scales with resolution. The saving is larger than the vendor framing implies.
HD full render, end to end
20 seconds × $0.17. The Full HD equivalent is $5.80, and a full-length draft pass is $1.20. Those three numbers are the whole cost model for a single-shot social clip.
Put the draft mechanic into a realistic iteration loop and the case makes itself. Eight full-length draft passes plus one HD full render at 20 seconds costs $13.00 — eight at $1.20 plus $3.40. Nine HD full renders, which is what the same amount of exploration costs without drafts, comes to $30.60. That is $17.60 saved on a single shot, or 57% off the exploration budget, and the saving scales linearly with how many variants your process actually needs before someone signs off.
05 — The FieldNormalising per-second pricing onto a per-minute axis.
Comparing video models on price is harder than it should be because the industry quotes in two incompatible units. BFL prices per second by mode and resolution. Independent leaderboards price per minute of finished video — Artificial Analysis states its API pricing column reflects the cost to generate one minute of 1080p video on the model creator’s API at default settings. Nothing lines up until one side is converted.
The table below does that conversion. Every FLUX 3 Video row starts from a published per-second rate and multiplies by 60; every competitor row starts from a listed per-minute price and divides by 60. The direction of each conversion is marked, because a derived number should never be mistaken for a quoted one.
| Tier or model | Per second | Per finished minute | Independent Arena Elo |
|---|---|---|---|
| FLUX 3 Video · BFL list price · per-second quoted, per-minute derived | |||
| Text-to-video / image-to-video, HD 720p | $0.17 | $10.20 | Not yet scored |
| Text-to-video / image-to-video, Full HD 1080p | $0.29 | $17.40 | Not yet scored |
| Video continuation, HD 720p | $0.43 | $25.80 | Not yet scored |
| Video continuation, Full HD 1080p | $0.54 | $32.40 | Not yet scored |
| Draft Mode, text-to-video / image-to-video | $0.06 | $3.60 | Preview render |
| Draft Mode, video continuation | $0.12 | $7.20 | Preview render |
| FLUX 3 Video · OpenRouter provider rate · differs on continuation only | |||
| Video continuation, 720p | $0.41 | $24.60 | Not yet scored |
| Video continuation, 1080p | $0.53 | $31.80 | Not yet scored |
| The rest of the field · per-minute listed, per-second derived unless marked vendor | |||
| Gemini Omni Flash | $0.10 | $6.00 | 1244 |
| MiniMax H3, 2K | $0.13 (vendor) | $7.80 | 1238 |
| Seedance 2.0, 720p | $0.15 | $9.07 | 1224 |
| Seedance 2.5 | Not published | Not published | Not yet scored |
| Kling 3.0 Standard, 720p | $0.25 | $15.12 | 1101 |
| Kling 3.0 Pro, 1080p | $0.34 | $20.16 | 1111 |
| Veo 3.1 Lite | $0.08 | $4.80 | 1089 |
| Veo 3.1 Fast | $0.15 | $9.00 | 1091 |
| Veo 3.1 | $0.40 | $24.00 | 1093 |
Digital Applied calculation. FLUX 3 Video rows: per-second prices as published by Black Forest Labs and, for the two continuation rows in the second group, by OpenRouter; per-minute figures are our conversion at ×60. Competitor rows: per-minute figures as listed by Artificial Analysis; per-second figures are our conversion at ÷60, except MiniMax H3 where the per-second rate is the vendor’s own published figure and the per-minute column is our ×60 conversion of it. Elo values are Artificial Analysis Text-to-Video Arena scores at the time of writing. Artificial Analysis defines its pricing column as the cost of one minute of 1080p video at default settings, so cross-vendor rows are indicative rather than contractual — quote from the surface you bill through.
Cost per finished minute · FLUX 3 Video tiers vs the field
Digital Applied calculation · vendor + AA pricesThe like-for-like row against Artificial Analysis pricing is FLUX 3 Video at Full HD, $17.40 per finished minute, because that is the only tier that matches the 1080p basis the leaderboard uses. On that axis FLUX 3 Video comes in 27.5% below Veo 3.1 at $24.00 and 13.7% below Kling 3.0 Pro at $20.16 — but 2.2× above MiniMax H3 at $7.80 and 2.9× above Gemini Omni Flash at $6.00, which happens to be the model sitting at the top of the text-to-video Arena.
Read that honestly and the positioning is mid-market, not price-leading. FLUX 3 Video undercuts the incumbent premium tier and sits well above the aggressive Chinese and Google-efficiency tiers. The exception is Draft Mode: at $3.60 per finished minute it is 25% below the cheapest scored row in the table, Veo 3.1 Lite at $4.80. That puts BFL’s cheapest published rate on iteration rather than on final renders.
For the wider economics of per-second video pricing and where these rates sit against the cost of the ad creative they replace, we unpacked the trend line in the per-second video pricing war, and the Kling side of the table is covered in our Kling 3 guide.
06 — Evidence QualityWhat the quality claims actually establish.
Every video launch arrives with a preference chart, and this one is no exception. What matters is reading the vendor’s own wording rather than the headline someone writes on top of it.
Three things follow from that sentence. First, it is an internal evaluation, so the rater pool, prompt set and comparison conditions are the vendor’s own. Second, on image-to-video the claim is a tie with Seedance 2.0 and a win over everyone else — not a win over Seedance. A great deal of launch coverage will round that up, and it is wrong to. Third, the comparison names Seedance 2.0, not Seedance 2.5, which is a different and newer model.
BFL published a rater-preference chart alongside the claim, but its headline figure appears in the chart caption rather than the body text, and it is a score from the vendor’s own evaluation rather than a public arena. We are not printing that number as a benchmark, for the same reason we would not print an internally-scored exam grade as an external qualification.
The independent picture is simply absent. At the time of writing, FLUX 3 Video had no row on either the text-to-video or image-to-video Artificial Analysis Arena leaderboard. Neither did Seedance 2.5. Both are too recent for a voting-based arena to have accumulated comparisons, which is a limitation of the method rather than a judgement on either model. For reference, the top of the text-to-video-with-audio board at the time of writing was Gemini Omni Flash at 1244 Elo, with MiniMax H3 second at 1238 and Seedance 2.0 720p third at 1224 — that is the band a new entrant needs to reach to be competitive on measured preference rather than vendor preference.
One piece of pre-release diligence is worth noting because it is rarely disclosed at all: BFL states it worked with third-party partner Cinder to evaluate FLUX 3 Video across supported modalities before release, including for non-consensual intimate imagery and CSAM risk. That is a process disclosure rather than a published result, but for agencies with brand-safety obligations, a named external evaluator is more than most video launches offer.
Our pre-GA snapshot of this field — written before H3 shipped and before Seedance 2.5 launched — is at the earlier video field comparison. This post supersedes its pricing and availability rows.
07 — Capability MatrixWhat each model can actually do.
Price is one axis; what the model will let you build is another. The matrix below covers the three models with a primary specification source retrieved for this piece — FLUX 3 Video, MiniMax H3 and Seedance 2.5. Cells that the primaries do not establish are marked as such rather than filled in from secondary coverage.
| Model | Clip length | Audio | Multi-shot in one generation | Continuation | Reference inputs |
|---|---|---|---|---|---|
| FLUX 3 Video | 5–20 s text-to-video and image-to-video; 5–15 s continuation | Generated alongside the frames — effects, ambience, lip-synced dialogue; on by default | Yes — multiple scenes and camera angles in one request | Yes — up to 4 s of existing video and audio in | 1–10 images, pinnable to exact timestamps |
| MiniMax H3 | 4–15 s | Native synced audio | Not established by the primary | Not established by the primary | Reference-to-video mode alongside text-to-video and image-to-video; instruction-based editing |
| Seedance 2.5 | 30 s single-pass one-take; multi-minute via multi-round continuation | Accepts audio clips as references; native audio generation not stated in the launch post | One-take single-pass generation | Yes — multi-round continuation | Up to 30 images, 10 videos and 10 audio clips |
| Kling 3.0 · Veo 3.1 | Neither vendor shipped a new numbered release in this window — Kling 3.0 dates to February 2026 and Veo 3.1 to January 2026 per Artificial Analysis. Their pricing and measured Arena positions appear in the table above; capability rows here are restricted to models with a primary specification source retrieved for this piece. | ||||
Digital Applied compilation from vendor primaries. FLUX 3 Video from the Black Forest Labs launch post and API documentation; MiniMax H3 from MiniMax platform documentation; Seedance 2.5 from its launch post. Seedance 2.5 API access was listed as coming soon in its own announcement, with no vendor-published price, so it is unpriced in the pricing table above — any per-minute figure circulating for it belongs to Seedance 2.0, a separate and older model.
Two availability notes change how these rows should be read. MiniMax H3 launched at the end of July with a live pay-as-you-go API, and we covered it in the MiniMax H3 launch write-up. Seedance 2.5 launched into consumer surfaces with API access described as coming soon in its own announcement, covered in the Seedance 2.5 one-take launch. A 30-second one-take capability you cannot yet call from an API is not, today, a pipeline option — which is the practical reason FLUX 3 Video and H3 are the two rows that matter for anyone building this quarter.
08 — Working MethodDraft first, render once.
The routing question is not “is FLUX 3 Video the best model” — with no independent score, that question has no defensible answer yet. The useful question is which jobs its specific shape fits, and where something cheaper or longer-running fits better.
Social clips that need sound
Lip-synced multilingual dialogue generated with the frames removes a whole post step, and the 9:16 and 1:1 ratios are native. Draft at $0.06/s until the shot is right, then enhance once at $0.17/s HD.
Hundreds of SKUs, short clips
Unit cost dominates. A 6-second HD clip is roughly $1.02 at list, and keyframe pinning keeps the product consistent across a set. Budget the draft passes explicitly — they are the line item that grows with SKU count.
Where FLUX 3 Video is not the answer
At $17.40 per finished 1080p minute it is 2.2× MiniMax H3 and 2.9× Gemini Omni Flash, both of which carry measured Arena scores it does not yet have. For volume work where preference data matters more than novelty, start there.
Beyond 20 seconds
Chain continuation passes — but price them properly: continuation runs $0.43/s HD on vendor list, more than double a fresh HD render, and caps at 15 seconds per pass. Seedance 2.5 targets this shape natively, once its API opens.
The workflow that falls out of the pricing is unusually explicit for a video model: iterate in Draft Mode, approve a shot, then use draft-enhance so the full render reproduces the approved subjects, composition and motion from the draft cache rather than reinterpreting the prompt. That is the difference between paying for exploration once and paying for it every time the model reinterprets a brief. Teams that skip drafts because the per-second rate looks cheap will spend multiples of what a draft-first process costs.
Where this heads next is worth stating as a projection rather than a fact. The pattern across this year’s video launches is that the headline specification — clip length, resolution, native audio — converges quickly, while the pricing structure is where vendors differentiate. BFL has priced a deep preview tier into this release, and the plausible next move is for iteration cost to become a pricing battleground in its own right, because it is what actually determines whether AI video is viable at catalogue or campaign scale. Independent Arena scores for both FLUX 3 Video and Seedance 2.5 should appear within weeks; until they do, treat every quality ranking in circulation, including the vendor’s, as provisional.
If you are pricing this into a campaign rather than a prototype, the catalogue-scale economics get their own treatment in our guide to product video at catalogue scale, and the media-side question of what generated video actually does to creative performance sits inside our paid media engagements. For teams building the generation pipeline itself rather than buying the output, our ecommerce practice runs the SKU-consistency and cost-modelling side.
09 — ConclusionA real release with an honest gap in the evidence.
Judge it on the price sheet, because the scoreboard is empty.
FLUX 3 Video’s general availability is a substantial release: 20 seconds, audio generated with the frames, multi-shot sequencing in a single request, timestamp-pinnable keyframes, and a continuation mode that carries dialogue across a seam. Those are capabilities you can verify against documentation and test on your own prompts today.
What you cannot verify yet is quality relative to the field. The comparison BFL published is an internal rater evaluation, and its image-to-video claim is a tie with Seedance 2.0 rather than a win. No independent arena had scored the model at the time of writing. That is not a criticism of the release — it is a day old — but it does mean any confident ranking you read this week is built on the vendor’s own data.
The durable insight from this launch is not about FLUX at all. It is that per-second pricing has become the marketing surface for video models the way per-token pricing became one for language models, and the same distortions have arrived with it: one quoted rate standing in for six, provider surfaces quietly disagreeing on the same model, and a preview tier that changes the real economics more than the headline rate does. Price the mode you will actually run, at the resolution you will actually ship, on the surface you will actually bill through. Everything else is a press release.