AI video spec sheets lead with two numbers — clip length and resolution — and both are routinely quoted as if they measured one thing. They don’t. “40 seconds” can mean one continuous generation pass or four chained extension rounds. “4K” can mean the pixels a model actually renders, a separate upscale pass, or a third mechanism that is neither. Where a vendor names which mechanism applies, it does so on its own page; most coverage flattens the result into a single headline figure.
The distinction is not pedantry. A 40-second cumulative ceiling reached through extensions behaves differently from a 30-second single pass: every extension introduces a seam, and whether a character’s face survives that seam depends on a number most vendors don’t publish at all — how much prior video the model reads before generating the next segment. Likewise, footage upscaled to 4K and footage rendered at a native resolution carry different detail, different artifacts, and in at least one documented case different surface availability.
This reference normalizes both spec numbers into one taxonomy — native single pass, chained extension with a stated context window, and upscale or regeneration — then scores five 2026 video models against it using each vendor’s own wording: Gemini Omni 1.1 Flash, FLUX 3 Video, Seedance 2.5, MiniMax H3, and Veo 3.1.
- 01Both spec numbers collapse distinct mechanisms.Duration collapses two: native single-pass generation and chained multi-round extension. Resolution collapses three: native output, post-hoc upscale, and in-context regeneration. Where a vendor names which one applies, it is on the vendor's own page; coverage usually doesn't name it at all.
- 02The biggest duration number here is not a single pass.Gemini Omni 1.1 Flash's 40 seconds is a cumulative ceiling reached in 10-second extension increments. The largest single-pass claim in this set is Seedance 2.5's 30 seconds, in ByteDance's own words a “single pass.”
- 03The seam variable is the prior-context window.Gemini Omni 1.1 Flash reads up to 10 seconds of prior video per extension — up from the final second only in previous models. FLUX 3's continuation mode reads up to 4 seconds. ByteDance states no figure for Seedance 2.5.
- 04No vendor in this set claims native 4K.Google describes both Omni 1.1 Flash's and Veo 3.1's 1080p/4K outputs as upscaling, BFL states FLUX 3 is native 720p with 1080p via upscaling, ByteDance states no resolution for Seedance 2.5, and MiniMax's native figure for H3 is 2K.
- 05Absence comes in three kinds — treat them differently.“Vendor does not state” (the page exists and omits the number), “page not located,” and “fetch method could not read it” are three different evidentiary situations. A spec table that doesn't distinguish them manufactures certainty.
01 — The CollapseTwo numbers, five mechanisms.
Read the vendor pages behind the 2026 video models and a pattern emerges: the headline spec numbers are honest, but they are answers to different questions. On the duration side there are two mechanisms — a native single pass, where the model generates the whole clip in one continuous run, and chained extension, where the model repeatedly appends new segments to existing output. On the resolution side there are three — native generation, post-hoc upscaling, and a third path MiniMax introduced that fits neither bucket. Two plus three: five distinct mechanisms compressed into two spec-sheet numbers.
This post is deliberately about one axis: the unit behind the number — what mechanism a duration or resolution claim actually describes. It is a different axis from availability — whether you can call the model at all — which we scored separately in the buildability audit of what actually ships. A model can be fully buildable with a misread spec, or correctly specced and unavailable; the two failure modes are independent, and the vendor status words quoted in the tables below are there to keep them separate.
02 — Duration ClaimsOne pass, or chained?
Start with the biggest number in the set. Google’s launch post for Gemini Omni 1.1 Flash, published August 27, 2026, states that you can extend videos in 10-second increments up to a total cumulative length of 40 seconds. That is chained extension: the 40-second figure is a ceiling across multiple rounds, not the length of any single generation. The same post adds first/last-frame control — you supply the starting and ending frames of a shot and the model generates continuous motion between them — and a 360p draft mode that Google frames as up to 60% faster and at a third of the cost of the standard 720p resolution. Draft-tier economics are their own story; here the relevant fact is that the headline duration is cumulative.
"You can extend videos in 10-second increments up to a total cumulative length of 40 seconds."— Google, Gemini Omni 1.1 Flash launch post, August 27, 2026
Now the single-pass claims. ByteDance’s own launch post for Seedance 2.5 states the model “can generate high-quality, 30-second audio-video clips in a single pass” — the vendor’s term is “single pass,” and “one-take” is media shorthand, not ByteDance’s verbatim wording. That doubles the 15 seconds ByteDance references for its predecessor, Seedance 2.0. Beyond the single pass, Seedance 2.5 also supports multi-round extension: users can “smoothly append subsequent shots to existing video outputs,” referencing a prior clip by name, up to “several minutes of content from one workflow.”
Black Forest Labs’ FLUX 3 Video, per its GA rollout, generates clips up to 20 seconds long from a single prompt in text-to-video and image-to-video modes — a single-pass generation with native audio in the same pass. Its separate video continuation mode is capped lower: 5 to 15 seconds of new output per step. The headline 20-second figure applies only to the single-pass modes, not uniformly across the product. And MiniMax H3 generates video with native stereo sound up to 15 seconds — with no extension mode stated in the announcement, and no statement of whether that length is one pass.
Headline duration claim · seconds · mechanism noted per row
Sources: blog.google · seed.bytedance.com · bfl.ai · minimax.io (each model's launch post)Lay the four claims side by side and the trap is visible: the largest headline number in the set — 40 seconds — is the only one that is not a single pass. Ranked by continuous generation, Seedance 2.5’s 30 seconds leads, FLUX 3’s 20 follows, and Gemini Omni 1.1 Flash’s per-round unit is 10. Neither ranking is more “true” — they answer different production questions. A continuous take matters for unbroken camera moves; a high extension ceiling matters for total shot length. A spec comparison that mixes the two units is comparing a marathon time to a relay time.
03 — Prior ContextThe seam number aggregators skip.
If a duration claim means chained extension, the next question is the one aggregators almost never ask: how much of the existing video does the model read before generating the next segment? That prior-context window is the mechanism variable behind seam quality. A model that conditions on one final frame can match momentary appearance but has no memory of how a character moved, what left the frame, or where the light was two shots ago. A model that reads ten seconds carries motion, identity, and scene state across the cut.
Google publishes its figure, and frames it as the generational change: Gemini Omni 1.1 Flash’s extension step analyzes up to 10 seconds of prior context — in Google’s own words, a leap from previous models that only referenced the final second. BFL publishes one too: FLUX 3’s continuation mode conditions on up to four seconds of existing video — a materially smaller window than its 20-second single-pass headline might suggest. ByteDance, by contrast, describes Seedance 2.5’s appending behavior qualitatively — carrying character, environment, and pacing forward — but states no numeric window at all.
Prior-context window per extension step · seconds of existing video read
Sources: blog.google (Omni 1.1 Flash launch) · bfl.ai (FLUX 3 Video, Part 1) · seed.bytedance.com (Seedance 2.5 launch)Two of the three current models charted carry a vendor-stated figure; the Seedance row is an honest zero-length bar, because the correct entry is “vendor does not state,” not a guess. The window sizes matter at the seam, a problem we’ve examined in the single-shot-vs-multi-shot coherence problem around the earlier Seedance generation.
04 — Duration TableThe duration taxonomy, four models scored.
The table below isolates the column that existing comparisons skip: the prior-context window per extension step.
| Model · source | Headline duration claim | Mechanism | Prior context per extension step | Vendor’s own status word |
|---|---|---|---|---|
| Gemini Omni 1.1 Flash · blog.google · 2026-08-27 | 40 seconds — cumulative ceiling, not one pass | Chained extension: “extend videos in 10-second increments” | “analyzes up to 10 seconds of prior context” — previous models “only referenced the final second” | “production-ready,” “rolling out across the Google developer ecosystem” |
| FLUX 3 Video · bfl.ai · 2026-08-04 | 20 seconds — single prompt, text- or image-to-video | Native single pass; separate continuation mode adds 5–15s of new output per step | Continuation conditions on “up to four seconds of existing video” | “generally available via the BFL API and select partners” — a gated GA (VentureBeat: “limited release to start”) |
| Seedance 2.5 · seed.bytedance.com · 2026-07-31 | 30 seconds — “in a single pass” (vendor wording; “one-take” is media shorthand) | Native single pass + multi-round append (“@Video 1… Continue from the visuals and subjects”) up to “several minutes” | Vendor does not state — no numeric figure in the launch post | Live on Jimeng and Doubao Pro at launch; API via Volcano Engine/BytePlus ModelArk “coming soon” |
| MiniMax H3 · minimax.io · 2026-07-31 | “up to 15 seconds” — with native stereo sound | Not specified — the vendor states the length, not the mechanism | No extension mode stated in the announcement | Weights announced but not yet public at publication — plans “to open up the model weights in the coming days” |
Recounting the table: of the four rows, two carry a single-pass headline — ByteDance says “in a single pass” for Seedance 2.5 in so many words, while BFL describes FLUX 3 as producing clips of up to twenty seconds from a single prompt — one is a chained ceiling (Gemini Omni 1.1 Flash), and one states a length without naming a mechanism (MiniMax H3). Exactly two rows carry a vendor-stated prior-context figure — Gemini’s up-to-10 seconds and FLUX 3’s up-to-4 seconds.
05 — Resolution ClaimsThree paths to the big number.
Resolution claims split three ways, and vendors are more precise about which path applies than the coverage that quotes them. Google’s Omni 1.1 Flash workflow is draft-then-upscale: the model generates at 360p or 720p, and 1080p or 4K arrive through a separate resolution pass — which is why “Gemini Omni 1.1 Flash generates 4K video” is a misreading, and “offers 4K via upscale” is the accurate sentence. Google’s Veo 3.1 materials use the same verb repeatedly: “Generate videos 1080p and 4K with state-of-the-art upscaling.” BFL is just as explicit for FLUX 3 Video: output “is released at HD, meaning 720p natively, with Full HD available through upscaling.”
Native generation
The resolution of the generation pass itself. FLUX 3 Video is native 720p; Gemini Omni 1.1 Flash drafts at 360p or generates standard 720p; MiniMax states H3 outputs up to 15 seconds at 2K — the only native figure above 1080p any vendor in this set claims.
Post-hoc upscale
A distinct step that raises resolution after the fact. Google's verb for Veo 3.1's 1080p/4K is “upscaling” in both posts we cite; Omni 1.1 Flash's 1080p/4K outputs are a separate resolution pass after the draft; BFL reaches Full HD “through upscaling.”
In-context regeneration
MiniMax's own term for how H3 reaches higher resolution: the H3 base model regenerates its own low-resolution output in-context. Explicitly not a traditional super-resolution upscaler, and not native single-pass generation at the higher resolution — a genuine third category most taxonomies have no slot for.
The third path deserves the emphasis. Upscalers are separate models bolted after generation; native generation is the diffusion pass itself. MiniMax’s in-context regeneration is neither: the same base model takes its own low-resolution output as context and regenerates it. Whether that produces upscaler-like artifacts or generation-like detail is an empirical question the announcement doesn’t settle — but collapsing it into “upscaling,” as most aggregator writeups do, erases the one mechanistic novelty in the set.
06 — Resolution TableThe resolution taxonomy, five models scored.
Same discipline as the duration table: the vendor’s verb for the higher-resolution path where one exists, with absences named rather than filled. The Seedance 2.5 row is the instructive one — a major launch whose own post contains no resolution specification at all.
| Model · source | Native output resolution | Higher-resolution path | Vendor’s own wording |
|---|---|---|---|
| Gemini Omni 1.1 Flash · blog.google | 360p draft · 720p standard | Upscale — a separate resolution pass after draft generation, to 1080p or 4K | “Generate polished, high-resolution 1080p or 4K outputs” — offered as a distinct step in the draft-then-upscale workflow |
| FLUX 3 Video · bfl.ai | 720p (HD) | Upscaling pass to 1080p (Full HD) | “released at HD, meaning 720p natively, with Full HD available through upscaling” |
| Seedance 2.5 · seed.bytedance.com | Not specified — vendor does not state | Not specified — vendor does not state | The launch post contains no resolution specification; pre-launch “4K” figures in circulation were an expected capability, not a confirmed spec |
| MiniMax H3 · minimax.io | 2K — vendor-stated native output | In-context regeneration — not a classic upscaler | “up to 15 seconds at 2K resolution”; the base model “regenerate[s] its own low-resolution output in-context” |
| Veo 3.1 · blog.google · cloud.google.com | Vendor does not state — neither Google post cited names a native generation resolution | Upscaling to 1080p and 4K; Google’s footnote restricts it to Flow, the Gemini API, and Vertex AI | “Generate videos 1080p and 4K with state-of-the-art upscaling”; “a new Veo upscaling capability” |
Recounting this table: five rows, of which three carry a vendor-stated native figure (Gemini Omni 1.1 Flash, FLUX 3, MiniMax H3) and two are absences of the first kind — Seedance 2.5 and Veo 3.1, where the pages cited simply do not state a native number. No row claims native 4K. And note the Veo footnote: Google’s January 2026 post restricts the 1080p/4K upscale to Flow, the Gemini API, and Vertex AI — explicitly excluding the Gemini app and YouTube surfaces.
Seedance 2.5 · combined inputs
Reference-input capacity at launch — images, video, audio, and 3D white models combined — up from 12 in Seedance 2.0, per launch-period coverage. A capability claim, not a duration or resolution claim; included here as context for what the launch post does specify while staying silent on resolution.
MiniMax H3 · vendor-stated
The only vendor-stated native resolution above 1080p in this five-model set. Reached higher through in-context regeneration rather than an upscaler — and, at the announcement's publication, with weights announced for release “in the coming days” rather than already public.
07 — Both DirectionsPress error runs both ways.
The definitional collapse is usually framed as hype deflation — vendors overstate, careful readers correct downward. The record around these five models shows something more interesting: third-party pages overstate vendors’ own numbers upward, in both of the directions this post tracks.
Case one: MiniMax H3. MiniMax’s own announcement states 2K. Yet multiple secondary pages advertise the model as native 4K at 60FPS. Case two: Veo 3.1. Google’s own posts say “upscaling,” repeatedly and in both the consumer and cloud framings — yet review-style aggregator pages claim the model achieves native 4K rather than upscaling, directly contradicting the vendor. In both cases the primary source is more modest than the coverage, and in both cases the correct citation is the vendor page, not the page with the bigger number.
This is the same failure mode we documented in the availability-language audit — coverage upgrading “preview” to “GA” — now operating on capability numbers. The mechanism is identical: speculation published before a launch outranks the launch in search for a while, and nobody goes back. The practical consequence for anyone budgeting a video pipeline is that resolution claims sourced from anywhere other than the vendor’s page carry a documented upward bias in this category.
08 — Field GuideHow to read the next spec sheet.
The taxonomy compresses into four questions. They take minutes to answer against a vendor page and will usually change which model a production plan actually needs.
One pass, or chained?
Find the increment. If the headline length is reached “in N-second increments” or by appending shots, it's a cumulative ceiling with seams — compare it against other ceilings, not against single-pass claims. Gemini Omni 1.1 Flash's 40s and Seedance 2.5's “several minutes” are ceilings; Seedance's 30s and FLUX 3's 20s are passes; MiniMax states H3's 15s without naming a mechanism.
What is the prior-context window?
The seconds of existing video read per extension step is the number that predicts identity survival at seams. In this set only Google (up to 10s) and BFL (up to 4s, continuation mode) publish one. If it isn't published, treat seam behavior as untested rather than inferring a figure.
Native, upscale, or regeneration?
Find the vendor's verb. “Upscaling” (Google for Veo, BFL for FLUX 3), a separate resolution pass (Omni 1.1 Flash), and “in-context regeneration” (MiniMax) are three different mechanisms with different artifact profiles. Then check the footnotes for surface restrictions — Veo's 1080p/4K upscale is restricted to Flow, the Gemini API, and Vertex AI.
Which kind of missing?
“Vendor does not state” means the page exists and omits the number — Seedance 2.5's resolution is this kind. “Page not located” means no page was found to check. “Fetch method could not read it” means a page exists but couldn't be retrieved. Only the first licenses the sentence “the vendor publishes no figure”; the other two license nothing.
Two implications follow from the pattern. First, on interpretation: the vendors are converging on a generate-low, deliver-high architecture — cheap native passes (360p drafts on Omni 1.1 Flash, native 720p on FLUX 3) with resolution added by a separate, separately-controlled step. That split exists because it matches how production actually works: iterate cheap, finish expensive. The spec-sheet confusion is a side effect of an architecture decision that is, on its own terms, sensible. Duration is following the same shape — short native passes, length added by extension — which is why the unpublished prior-context window is becoming the real differentiator while headline seconds converge.
Second, projecting forward: as extension ceilings stretch, expect vendors that publish their context window to make it a marketing number — Google already frames 10-versus-1 seconds as the generational leap — and expect the gap between ceiling and pass to widen faster than single-pass length improves, because chaining is the cheaper axis. When a spec number determines real budget — and per-second pricing differs by mechanism, as we break down in translating these units into cost per finished second — the mechanism check above is the difference between a plan and a surprise. Mapping claims like these to production workflows is exactly the kind of evaluation our AI transformation engagements run before a stack decision, and the same discipline drives how our content engine treats vendor claims in published work: the vendor’s page, the vendor’s verb, the vendor’s number.
09 — ConclusionQuote the mechanism, not the number.
A spec number without its mechanism is not yet a fact.
The five models in this reference are not guilty of vague specs — the opposite. Google says “10-second increments” and “upscaling.” BFL says “720p natively.” ByteDance says “single pass” and declines to name a resolution. MiniMax coins in-context regeneration for a mechanism that needed a new name. The precision exists at the source; it is lost in transmission.
The working rules are short. A duration claim is a pass length or a ceiling — establish which before comparing. An extension claim is only as good as its prior-context window, and most vendors don’t publish one. A resolution claim is native, upscaled, or regenerated — the vendor’s verb tells you which. And an absent number is one of three different absences, only one of which says anything about the vendor.
The satisfying part of this corner of the industry is that the primary sources are short, public, and unusually precise. Reading them slightly more carefully than the headlines did is the cheapest quality bar any team shipping AI-video claims, or AI-video budgets, can adopt today.