AI DevelopmentDecision Matrix13 min readPublished August 27, 2026

Five 2026 video models · 2+3 mechanisms behind two spec numbers · the prior-context figure aggregators skip

What 40 Seconds and 4K actually mean in AI video

Gemini Omni 1.1 Flash’s 40 seconds is a cumulative extension ceiling, not one generation pass, and its 4K is an upscale step. FLUX 3 Video’s 1080p is upscaled from native 720p. Seedance 2.5’s launch post names no resolution at all. The outlier is MiniMax H3: vendor-stated native 2K, raised by a mechanism MiniMax calls in-context regeneration. This is the reference for what each number actually claims.

DA
Digital Applied Team
Senior strategists · Published Aug 27, 2026
PublishedAug 27, 2026
Read time13 min
SourcesPrimary + secondary
Gemini Omni 1.1 Flash ceiling
40s
cumulative, via chained 10s extensions
not one pass
Prior context per extension
≤10s
vendor-stated; was final second only
MiniMax H3 native output
2K
vendor-stated; some pages claim 4K
the outlier
Seedance 2.5 resolution figures
0
stated in its own launch post

AI video spec sheets lead with two numbers — clip length and resolution — and both are routinely quoted as if they measured one thing. They don’t. “40 seconds” can mean one continuous generation pass or four chained extension rounds. “4K” can mean the pixels a model actually renders, a separate upscale pass, or a third mechanism that is neither. Where a vendor names which mechanism applies, it does so on its own page; most coverage flattens the result into a single headline figure.

The distinction is not pedantry. A 40-second cumulative ceiling reached through extensions behaves differently from a 30-second single pass: every extension introduces a seam, and whether a character’s face survives that seam depends on a number most vendors don’t publish at all — how much prior video the model reads before generating the next segment. Likewise, footage upscaled to 4K and footage rendered at a native resolution carry different detail, different artifacts, and in at least one documented case different surface availability.

This reference normalizes both spec numbers into one taxonomy — native single pass, chained extension with a stated context window, and upscale or regeneration — then scores five 2026 video models against it using each vendor’s own wording: Gemini Omni 1.1 Flash, FLUX 3 Video, Seedance 2.5, MiniMax H3, and Veo 3.1.

Key takeaways
  1. 01
    Both spec numbers collapse distinct mechanisms.Duration collapses two: native single-pass generation and chained multi-round extension. Resolution collapses three: native output, post-hoc upscale, and in-context regeneration. Where a vendor names which one applies, it is on the vendor's own page; coverage usually doesn't name it at all.
  2. 02
    The biggest duration number here is not a single pass.Gemini Omni 1.1 Flash's 40 seconds is a cumulative ceiling reached in 10-second extension increments. The largest single-pass claim in this set is Seedance 2.5's 30 seconds, in ByteDance's own words a “single pass.”
  3. 03
    The seam variable is the prior-context window.Gemini Omni 1.1 Flash reads up to 10 seconds of prior video per extension — up from the final second only in previous models. FLUX 3's continuation mode reads up to 4 seconds. ByteDance states no figure for Seedance 2.5.
  4. 04
    No vendor in this set claims native 4K.Google describes both Omni 1.1 Flash's and Veo 3.1's 1080p/4K outputs as upscaling, BFL states FLUX 3 is native 720p with 1080p via upscaling, ByteDance states no resolution for Seedance 2.5, and MiniMax's native figure for H3 is 2K.
  5. 05
    Absence comes in three kinds — treat them differently.“Vendor does not state” (the page exists and omits the number), “page not located,” and “fetch method could not read it” are three different evidentiary situations. A spec table that doesn't distinguish them manufactures certainty.

01The CollapseTwo numbers, five mechanisms.

Read the vendor pages behind the 2026 video models and a pattern emerges: the headline spec numbers are honest, but they are answers to different questions. On the duration side there are two mechanisms — a native single pass, where the model generates the whole clip in one continuous run, and chained extension, where the model repeatedly appends new segments to existing output. On the resolution side there are three — native generation, post-hoc upscaling, and a third path MiniMax introduced that fits neither bucket. Two plus three: five distinct mechanisms compressed into two spec-sheet numbers.

This post is deliberately about one axis: the unit behind the number — what mechanism a duration or resolution claim actually describes. It is a different axis from availability — whether you can call the model at all — which we scored separately in the buildability audit of what actually ships. A model can be fully buildable with a misread spec, or correctly specced and unavailable; the two failure modes are independent, and the vendor status words quoted in the tables below are there to keep them separate.

02Duration ClaimsOne pass, or chained?

Start with the biggest number in the set. Google’s launch post for Gemini Omni 1.1 Flash, published August 27, 2026, states that you can extend videos in 10-second increments up to a total cumulative length of 40 seconds. That is chained extension: the 40-second figure is a ceiling across multiple rounds, not the length of any single generation. The same post adds first/last-frame control — you supply the starting and ending frames of a shot and the model generates continuous motion between them — and a 360p draft mode that Google frames as up to 60% faster and at a third of the cost of the standard 720p resolution. Draft-tier economics are their own story; here the relevant fact is that the headline duration is cumulative.

"You can extend videos in 10-second increments up to a total cumulative length of 40 seconds."— Google, Gemini Omni 1.1 Flash launch post, August 27, 2026

Now the single-pass claims. ByteDance’s own launch post for Seedance 2.5 states the model “can generate high-quality, 30-second audio-video clips in a single pass” — the vendor’s term is “single pass,” and “one-take” is media shorthand, not ByteDance’s verbatim wording. That doubles the 15 seconds ByteDance references for its predecessor, Seedance 2.0. Beyond the single pass, Seedance 2.5 also supports multi-round extension: users can “smoothly append subsequent shots to existing video outputs,” referencing a prior clip by name, up to “several minutes of content from one workflow.”

Black Forest Labs’ FLUX 3 Video, per its GA rollout, generates clips up to 20 seconds long from a single prompt in text-to-video and image-to-video modes — a single-pass generation with native audio in the same pass. Its separate video continuation mode is capped lower: 5 to 15 seconds of new output per step. The headline 20-second figure applies only to the single-pass modes, not uniformly across the product. And MiniMax H3 generates video with native stereo sound up to 15 seconds — with no extension mode stated in the announcement, and no statement of whether that length is one pass.

Headline duration claim · seconds · mechanism noted per row

Sources: blog.google · seed.bytedance.com · bfl.ai · minimax.io (each model's launch post)
Gemini Omni 1.1 Flashcumulative ceiling — chained 10s extensions, not one pass
40s
Seedance 2.5single pass (vendor wording) · extension: several minutes
30s
FLUX 3 Videosingle pass, text/image-to-video · continuation adds 5–15s
20s
MiniMax H3mechanism not stated · no extension mode stated
15s

Lay the four claims side by side and the trap is visible: the largest headline number in the set — 40 seconds — is the only one that is not a single pass. Ranked by continuous generation, Seedance 2.5’s 30 seconds leads, FLUX 3’s 20 follows, and Gemini Omni 1.1 Flash’s per-round unit is 10. Neither ranking is more “true” — they answer different production questions. A continuous take matters for unbroken camera moves; a high extension ceiling matters for total shot length. A spec comparison that mixes the two units is comparing a marathon time to a relay time.

03Prior ContextThe seam number aggregators skip.

If a duration claim means chained extension, the next question is the one aggregators almost never ask: how much of the existing video does the model read before generating the next segment? That prior-context window is the mechanism variable behind seam quality. A model that conditions on one final frame can match momentary appearance but has no memory of how a character moved, what left the frame, or where the light was two shots ago. A model that reads ten seconds carries motion, identity, and scene state across the cut.

Google publishes its figure, and frames it as the generational change: Gemini Omni 1.1 Flash’s extension step analyzes up to 10 seconds of prior context — in Google’s own words, a leap from previous models that only referenced the final second. BFL publishes one too: FLUX 3’s continuation mode conditions on up to four seconds of existing video — a materially smaller window than its 20-second single-pass headline might suggest. ByteDance, by contrast, describes Seedance 2.5’s appending behavior qualitatively — carrying character, environment, and pacing forward — but states no numeric window at all.

Prior-context window per extension step · seconds of existing video read

Sources: blog.google (Omni 1.1 Flash launch) · bfl.ai (FLUX 3 Video, Part 1) · seed.bytedance.com (Seedance 2.5 launch)
Gemini Omni 1.1 Flash“analyzes up to 10 seconds of prior context”
≤10s
FLUX 3 Video · continuationconditions on “up to four seconds of existing video”
≤4s
Previous Omni models“only referenced the final second” — Google's own contrast
1s
Seedance 2.5 · extensionvendor does not state a numeric figure

Two of the three current models charted carry a vendor-stated figure; the Seedance row is an honest zero-length bar, because the correct entry is “vendor does not state,” not a guess. The window sizes matter at the seam, a problem we’ve examined in the single-shot-vs-multi-shot coherence problem around the earlier Seedance generation.

Why 10 seconds vs 1 second is the story
Google’s own framing places the mechanism change above the headline ceiling: the leap is not 40 seconds of total length but ten seconds of memory at every seam where previous models carried one. When you evaluate any extension-based duration claim, ask for this number first. If the vendor doesn’t publish it — and in this five-model set, most don’t — treat seam behavior as untested, not as implied by the headline figure.

04Duration TableThe duration taxonomy, four models scored.

The table below isolates the column that existing comparisons skip: the prior-context window per extension step.

Duration claims across four 2026 AI video models: the headline number, the mechanism behind it, the prior-context window per extension step, and the vendor’s own status language. As of August 27, 2026; vendor pages can change.
Model · sourceHeadline duration claimMechanismPrior context per extension stepVendor’s own status word
Gemini Omni 1.1 Flash · blog.google · 2026-08-2740 seconds — cumulative ceiling, not one passChained extension: “extend videos in 10-second increments”“analyzes up to 10 seconds of prior context” — previous models “only referenced the final second”“production-ready,” “rolling out across the Google developer ecosystem”
FLUX 3 Video · bfl.ai · 2026-08-0420 seconds — single prompt, text- or image-to-videoNative single pass; separate continuation mode adds 5–15s of new output per stepContinuation conditions on “up to four seconds of existing video”“generally available via the BFL API and select partners” — a gated GA (VentureBeat: “limited release to start”)
Seedance 2.5 · seed.bytedance.com · 2026-07-3130 seconds — “in a single pass” (vendor wording; “one-take” is media shorthand)Native single pass + multi-round append (“@Video 1… Continue from the visuals and subjects”) up to “several minutes”Vendor does not state — no numeric figure in the launch postLive on Jimeng and Doubao Pro at launch; API via Volcano Engine/BytePlus ModelArk “coming soon”
MiniMax H3 · minimax.io · 2026-07-31“up to 15 seconds” — with native stereo soundNot specified — the vendor states the length, not the mechanismNo extension mode stated in the announcementWeights announced but not yet public at publication — plans “to open up the model weights in the coming days”

Recounting the table: of the four rows, two carry a single-pass headline — ByteDance says “in a single pass” for Seedance 2.5 in so many words, while BFL describes FLUX 3 as producing clips of up to twenty seconds from a single prompt — one is a chained ceiling (Gemini Omni 1.1 Flash), and one states a length without naming a mechanism (MiniMax H3). Exactly two rows carry a vendor-stated prior-context figure — Gemini’s up-to-10 seconds and FLUX 3’s up-to-4 seconds.

05Resolution ClaimsThree paths to the big number.

Resolution claims split three ways, and vendors are more precise about which path applies than the coverage that quotes them. Google’s Omni 1.1 Flash workflow is draft-then-upscale: the model generates at 360p or 720p, and 1080p or 4K arrive through a separate resolution pass — which is why “Gemini Omni 1.1 Flash generates 4K video” is a misreading, and “offers 4K via upscale” is the accurate sentence. Google’s Veo 3.1 materials use the same verb repeatedly: “Generate videos 1080p and 4K with state-of-the-art upscaling.” BFL is just as explicit for FLUX 3 Video: output “is released at HD, meaning 720p natively, with Full HD available through upscaling.”

Path 1
Native generation
the pixels the model actually renders

The resolution of the generation pass itself. FLUX 3 Video is native 720p; Gemini Omni 1.1 Flash drafts at 360p or generates standard 720p; MiniMax states H3 outputs up to 15 seconds at 2K — the only native figure above 1080p any vendor in this set claims.

FLUX 3 · Omni 1.1 Flash · MiniMax H3
Path 2
Post-hoc upscale
a separate pass after generation

A distinct step that raises resolution after the fact. Google's verb for Veo 3.1's 1080p/4K is “upscaling” in both posts we cite; Omni 1.1 Flash's 1080p/4K outputs are a separate resolution pass after the draft; BFL reaches Full HD “through upscaling.”

Veo 3.1 · Omni 1.1 Flash · FLUX 3
Path 3
In-context regeneration
the base model re-runs its own output

MiniMax's own term for how H3 reaches higher resolution: the H3 base model regenerates its own low-resolution output in-context. Explicitly not a traditional super-resolution upscaler, and not native single-pass generation at the higher resolution — a genuine third category most taxonomies have no slot for.

MiniMax H3 — unique in this set

The third path deserves the emphasis. Upscalers are separate models bolted after generation; native generation is the diffusion pass itself. MiniMax’s in-context regeneration is neither: the same base model takes its own low-resolution output as context and regenerates it. Whether that produces upscaler-like artifacts or generation-like detail is an empirical question the announcement doesn’t settle — but collapsing it into “upscaling,” as most aggregator writeups do, erases the one mechanistic novelty in the set.

06Resolution TableThe resolution taxonomy, five models scored.

Same discipline as the duration table: the vendor’s verb for the higher-resolution path where one exists, with absences named rather than filled. The Seedance 2.5 row is the instructive one — a major launch whose own post contains no resolution specification at all.

Resolution claims across five 2026 AI video models: native output, the path to the larger number, and the vendor’s own wording for that path. As of August 27, 2026; vendor pages can change.
Model · sourceNative output resolutionHigher-resolution pathVendor’s own wording
Gemini Omni 1.1 Flash · blog.google360p draft · 720p standardUpscale — a separate resolution pass after draft generation, to 1080p or 4K“Generate polished, high-resolution 1080p or 4K outputs” — offered as a distinct step in the draft-then-upscale workflow
FLUX 3 Video · bfl.ai720p (HD)Upscaling pass to 1080p (Full HD)“released at HD, meaning 720p natively, with Full HD available through upscaling”
Seedance 2.5 · seed.bytedance.comNot specified — vendor does not stateNot specified — vendor does not stateThe launch post contains no resolution specification; pre-launch “4K” figures in circulation were an expected capability, not a confirmed spec
MiniMax H3 · minimax.io2K — vendor-stated native outputIn-context regeneration — not a classic upscaler“up to 15 seconds at 2K resolution”; the base model “regenerate[s] its own low-resolution output in-context”
Veo 3.1 · blog.google · cloud.google.comVendor does not state — neither Google post cited names a native generation resolutionUpscaling to 1080p and 4K; Google’s footnote restricts it to Flow, the Gemini API, and Vertex AI“Generate videos 1080p and 4K with state-of-the-art upscaling”; “a new Veo upscaling capability”

Recounting this table: five rows, of which three carry a vendor-stated native figure (Gemini Omni 1.1 Flash, FLUX 3, MiniMax H3) and two are absences of the first kind — Seedance 2.5 and Veo 3.1, where the pages cited simply do not state a native number. No row claims native 4K. And note the Veo footnote: Google’s January 2026 post restricts the 1080p/4K upscale to Flow, the Gemini API, and Vertex AI — explicitly excluding the Gemini app and YouTube surfaces.

Reference load
Seedance 2.5 · combined inputs
50

Reference-input capacity at launch — images, video, audio, and 3D white models combined — up from 12 in Seedance 2.0, per launch-period coverage. A capability claim, not a duration or resolution claim; included here as context for what the launch post does specify while staying silent on resolution.

up from 12 in Seedance 2.0
The one native-2K row
MiniMax H3 · vendor-stated
2K

The only vendor-stated native resolution above 1080p in this five-model set. Reached higher through in-context regeneration rather than an upscaler — and, at the announcement's publication, with weights announced for release “in the coming days” rather than already public.

15s max

07Both DirectionsPress error runs both ways.

The definitional collapse is usually framed as hype deflation — vendors overstate, careful readers correct downward. The record around these five models shows something more interesting: third-party pages overstate vendors’ own numbers upward, in both of the directions this post tracks.

Case one: MiniMax H3. MiniMax’s own announcement states 2K. Yet multiple secondary pages advertise the model as native 4K at 60FPS. Case two: Veo 3.1. Google’s own posts say “upscaling,” repeatedly and in both the consumer and cloud framings — yet review-style aggregator pages claim the model achieves native 4K rather than upscaling, directly contradicting the vendor. In both cases the primary source is more modest than the coverage, and in both cases the correct citation is the vendor page, not the page with the bigger number.

Documented overstatement · MiniMax H3
A community writeup on Hugging Face flags the pattern directly: several pages advertise Hailuo 3.0 as a “native 4K 60FPS” model, but MiniMax’s own announcement says 2K — pages written speculatively before launch and never updated. The repair procedure is mechanical: when a spec number appears only on third-party pages and the vendor’s own page states a smaller one, the vendor’s number wins, and the delta is itself worth reporting.

This is the same failure mode we documented in the availability-language audit — coverage upgrading “preview” to “GA” — now operating on capability numbers. The mechanism is identical: speculation published before a launch outranks the launch in search for a while, and nobody goes back. The practical consequence for anyone budgeting a video pipeline is that resolution claims sourced from anywhere other than the vendor’s page carry a documented upward bias in this category.

08Field GuideHow to read the next spec sheet.

The taxonomy compresses into four questions. They take minutes to answer against a vendor page and will usually change which model a production plan actually needs.

Duration
One pass, or chained?

Find the increment. If the headline length is reached “in N-second increments” or by appending shots, it's a cumulative ceiling with seams — compare it against other ceilings, not against single-pass claims. Gemini Omni 1.1 Flash's 40s and Seedance 2.5's “several minutes” are ceilings; Seedance's 30s and FLUX 3's 20s are passes; MiniMax states H3's 15s without naming a mechanism.

Find the increment size
Seams
What is the prior-context window?

The seconds of existing video read per extension step is the number that predicts identity survival at seams. In this set only Google (up to 10s) and BFL (up to 4s, continuation mode) publish one. If it isn't published, treat seam behavior as untested rather than inferring a figure.

Ask for the seconds figure
Resolution
Native, upscale, or regeneration?

Find the vendor's verb. “Upscaling” (Google for Veo, BFL for FLUX 3), a separate resolution pass (Omni 1.1 Flash), and “in-context regeneration” (MiniMax) are three different mechanisms with different artifact profiles. Then check the footnotes for surface restrictions — Veo's 1080p/4K upscale is restricted to Flow, the Gemini API, and Vertex AI.

Find the vendor's verb
Absences
Which kind of missing?

“Vendor does not state” means the page exists and omits the number — Seedance 2.5's resolution is this kind. “Page not located” means no page was found to check. “Fetch method could not read it” means a page exists but couldn't be retrieved. Only the first licenses the sentence “the vendor publishes no figure”; the other two license nothing.

Name the absence type

Two implications follow from the pattern. First, on interpretation: the vendors are converging on a generate-low, deliver-high architecture — cheap native passes (360p drafts on Omni 1.1 Flash, native 720p on FLUX 3) with resolution added by a separate, separately-controlled step. That split exists because it matches how production actually works: iterate cheap, finish expensive. The spec-sheet confusion is a side effect of an architecture decision that is, on its own terms, sensible. Duration is following the same shape — short native passes, length added by extension — which is why the unpublished prior-context window is becoming the real differentiator while headline seconds converge.

Second, projecting forward: as extension ceilings stretch, expect vendors that publish their context window to make it a marketing number — Google already frames 10-versus-1 seconds as the generational leap — and expect the gap between ceiling and pass to widen faster than single-pass length improves, because chaining is the cheaper axis. When a spec number determines real budget — and per-second pricing differs by mechanism, as we break down in translating these units into cost per finished second — the mechanism check above is the difference between a plan and a surprise. Mapping claims like these to production workflows is exactly the kind of evaluation our AI transformation engagements run before a stack decision, and the same discipline drives how our content engine treats vendor claims in published work: the vendor’s page, the vendor’s verb, the vendor’s number.

09ConclusionQuote the mechanism, not the number.

The unit behind the number, August 2026

A spec number without its mechanism is not yet a fact.

The five models in this reference are not guilty of vague specs — the opposite. Google says “10-second increments” and “upscaling.” BFL says “720p natively.” ByteDance says “single pass” and declines to name a resolution. MiniMax coins in-context regeneration for a mechanism that needed a new name. The precision exists at the source; it is lost in transmission.

The working rules are short. A duration claim is a pass length or a ceiling — establish which before comparing. An extension claim is only as good as its prior-context window, and most vendors don’t publish one. A resolution claim is native, upscaled, or regenerated — the vendor’s verb tells you which. And an absent number is one of three different absences, only one of which says anything about the vendor.

The satisfying part of this corner of the industry is that the primary sources are short, public, and unusually precise. Reading them slightly more carefully than the headlines did is the cheapest quality bar any team shipping AI-video claims, or AI-video budgets, can adopt today.

Put vendor claims to work — accurately

The spec sheet is a starting point — the vendor page is the contract.

Our team evaluates video and creative AI models against your actual production workflow — mechanism-level spec verification, cost-per-finished-second modeling, and pilot builds, delivered in days not quarters.

Free consultationExpert guidanceTailored solutions
What we work on

AI video & creative model engagements

  • Model selection scored on vendor-verbatim specs
  • Cost modeling — drafts, upscales, extension rounds
  • Seam-quality testing for extension workflows
  • Pipeline pilots on Veo, FLUX, and API-first stacks
  • Claims review for published AI-video content
FAQ · AI video spec claims

The questions spec sheets don’t answer.

No. Google's launch post states you can extend videos in 10-second increments up to a total cumulative length of 40 seconds — the 40-second figure is a ceiling across chained extension rounds, not the length of a single generation. What makes the extensions usable is the prior-context window: each extension step analyzes up to 10 seconds of prior video, which Google contrasts with previous models that only referenced the final second. The model also supports first/last-frame control, generating continuous motion between two supplied keyframes. If your workflow needs a genuinely continuous take, compare single-pass claims instead — in this set the largest is Seedance 2.5's vendor-stated 30 seconds.
Related dispatches

Continue exploring AI video claims.