Gemini Omni 1.1 Flash arrived on August 27, 2026 with the two things developers kept asking Google’s video model for: scenes longer than a single pass, and a way to iterate without paying production rates. Extension now chains 10-second passes to a 40-second cumulative ceiling, and a 360p draft tier reportedly cuts the per-second price to roughly a third of the 720p rate.
The launch also carries two caveats that most same-day coverage buried. First, Google’s own API documentation labels the 1080p and 4K output options “upscaled” — the vendor itself never claims native 4K generation. Second, the model’s status splits by surface: generally available on the paid Gemini Developer API, yet listed as Preview on Google’s Enterprise Agent Platform under a different model ID, on the same release day.
This launch report covers what shipped, how the extension mechanic actually works, what the draft-then-upscale workflow means for quality, what the pricing page really says versus what the press reported, and which platform surface a team should build against. Every claim below traces to Google’s launch materials or named same-day coverage.
- 01Scenes now reach 40 seconds, ten at a time.Google states videos extend in 10-second increments to a 40-second cumulative total. Each generation pass remains short — XenoSpectrum puts individual outputs at 3 to 10 seconds — so 40 seconds is chained, never one-shot.
- 02Extension reads 10 seconds of prior context.Google’s post says the model can now analyze up to 10 seconds of prior footage, versus previous models that referenced only the final second. That context window, not the ceiling, is the continuity upgrade.
- 031080p and 4K are upscales — per Google’s own docs.The API resolution parameter lists four values, and the docs describe the top two verbatim as “1080p output (upscaled)” and “4k output (upscaled).” Native generation tops out below that; plan quality at capture resolution.
- 04Billing is tokens; the $/s table is reported, not printed.Google bills video output at 5,792 tokens per second of 720p, about $0.10/s. The four-tier $0.03/$0.10/$0.15/$0.30 per-second table comes from XenoSpectrum and The Decoder — consistent with Google’s token math, but not printed as a table on the pricing page.
- 05GA on the Gemini API, Preview on the enterprise surface.Google’s pricing page calls the model generally available on the paid Gemini API, while the Enterprise Agent Platform lists gemini-omni-1.1-flash-preview at Launch stage: Preview — same day, same vendor, different terms.
01 — What ShippedOne model update, five creative controls.
The announcement came via Google’s official blog, authored by Anish Nangia and Alisa Fortin, Product Managers at Google DeepMind. Their framing: “Today, we’re introducing Gemini Omni 1.1 Flash, a new suite of creative controls and generative video capabilities to support developers... today’s updates make Omni 1.1 production-ready for professional use via the Gemini API in Google AI Studio.” The post’s subhead promises “studio-quality video production, including the ability to extend a scene, first and last frame interpolation, crisp 4K upscaling, faster prototyping, and more” — note that Google’s own subhead already says “upscaling,” not “generation,” for the 4K capability.
Omni 1.1 is an update to the Gemini Omni Flash line that debuted at Google I/O in May — we covered the original Gemini Omni Flash launch and its image-generation sibling when they shipped. At that first launch, per XenoSpectrum’s historical framing, Google’s own documentation pointed developers to Veo 3.1 for scene extension and final-frame specification. Omni 1.1 absorbs both capabilities into Omni itself.
Scene extension
Chain extension passes onto an initial generation, with the model reading up to 10 seconds of prior context per pass. End-only: no mid-clip insertion, no backward extension.
Keyframe interpolation
Specify the starting and ending frames of a shot and the model generates continuous video between them — camera orbits, zoom transitions, looping clips.
360p draft mode
Lightweight previews Google says run up to 60% faster by system throughput and at a third of the cost of standard 720p. Draft cheap, upscale what survives review.
02 — Scene ExtensionHow 40 seconds actually gets built.
The mechanism matters more than the headline number. Google’s wording is precise: “You can extend videos in 10-second increments up to a total cumulative length of 40 seconds.” Cumulative is the operative word. XenoSpectrum’s same-day technical read — the best independent one in the coverage set — puts individual generation passes at 3 to 10 seconds at 24fps, with the 40-second total built by chaining extension passes onto an already-generated clip. A 40-second scene is a sequence of passes, never a single one.
What makes the chaining viable is the context window behind each pass. Per Google, the model now analyzes up to 10 seconds of prior footage when extending, where previous models referenced only the final second. Ten seconds of context means motion arcs, lighting, and subject identity have material signal to carry across the seam — one second meant the model was effectively guessing from a snapshot. For what duration and resolution claims across the video field actually mean as units, see our duration and resolution claims reference, published today.
Cumulative, not one-shot
Extension passes add 10 seconds at a time onto an initial 3-to-10-second generation. The ceiling is a chain limit, not a single-pass output length.
Prior footage analyzed
Each extension pass reads up to 10 seconds of prior context. Google’s post contrasts this with previous models that referenced only the final second.
Consistency inputs
Up to three reference videos of up to 3 seconds each carry character and motion consistency. Per XenoSpectrum’s read of the docs, cross-referencing across multiple videos is unsupported and stacking them may degrade results.
Two structural limits are easy to miss in the launch framing. Extension is end-only — there is no mid-clip insertion and no backward extension from the start of a clip, so edits inside an existing sequence still mean regenerating from the cut point. And externally uploaded videos used as an extension source must be 10 seconds or shorter. XenoSpectrum, reporting on Google’s terms, also notes that in the European Economic Area, Switzerland, and the UK, editing and extending externally uploaded videos face additional restrictions — teams in those regions should verify the current terms against their specific workflow before building on the upload path. On language, XenoSpectrum’s read of the documentation is that English is the only language described as fully supported, with other languages, including Japanese, stated as not evaluated.
"With Omni 1.1, the model can now analyze up to 10 seconds of prior context — a leap from previous models that only referenced the final second."— Google DeepMind, Omni 1.1 Flash launch post, August 27, 2026
03 — Creative ControlsKeyframes in, motion out — and drafts at a third of the cost.
The second control is first-and-last-frame interpolation. Google’s description: “Achieve smooth transitions and camera movements by specifying the starting and ending frames of a shot. Omni 1.1 generates continuous video between two keyframes, making it ideal for complex camera orbits, zoom transitions, or seamless looping clips.” For teams that storyboard in stills, this inverts the workflow — art-direct the two frames you care about, let the model solve the motion between them.
The third is the draft tier, and it is the one with budget implications. Google’s pitch: “Generate lightweight previews in 360p resolution up to 60% faster* and at a third of the cost compared to Omni 1.1’s standard 720p resolution.” That asterisk is Google’s own, and the footnote it points to matters: “*Up to 60% faster generation based on system throughput of 360p vs. 720p resolution.” A system-throughput claim is a statement about aggregate capacity, not a guarantee that any single generation lands 60% sooner — worth knowing before anyone promises turnaround times downstream.
Notably, per XenoSpectrum’s read of Google Flow’s guidance, the staged workflow is Google’s own recommendation: draft at 360p, review at 720p, then export to 1080p or 4K. The vendor is not positioning one-shot high-resolution generation as the intended path — iteration happens at draft resolution, and the expensive tiers exist to finish, not to explore.
04 — The Fine Print4K is an upscale — Google’s own docs say so.
This is the claim to get precisely right, because the primary source could not be clearer. The resolution parameter table in Google’s own developer documentation lists exactly four values: 360p (“360p output resolution”), 720p (“720p output resolution (default)”), 1080p — described verbatim as “1080p output (upscaled)” — and 4k, described verbatim as “4k output (upscaled).” That is the literal API parameter documentation, not marketing copy, and it is the strongest form of vendor confirmation there is: the two top tiers are upscales of lower-resolution generation, not native output.
The launch post is consistent with the docs. Its section header for the capability reads “Upscale up to 4K resolution,” with body copy promising to “Generate polished, high-resolution 1080p or 4K outputs that are ready for professional production with Omni 1.1.” Nowhere does Google claim native 4K generation — the “40-second 4K video in one shot” framing that circulated in some aggregator coverage is a misreading on both axes, since the 40 seconds is chained and the 4K is upscaled. XenoSpectrum built its entire same-day analysis around correcting exactly that misreading.
The practical consequence for production teams: detail that was not captured at generation resolution will not be conjured by the upscale step. Quality decisions get made where generation happens, which by Google’s own recommended workflow is 360p drafting and 720p review. Treat the 1080p and 4K tiers as delivery formats, and evaluate whether upscaled output clears your client’s bar on your own footage before promising 4K deliverables.
05 — PricingTokens are the mechanism; $/s is the shorthand.
Google does not bill this model by the second — it bills by token, and the per-second figures everyone quotes are derived. The pricing page’s own words: “Billing is based on total output token consumption, calculated at a rate of 5,792 tokens per second of 720p video. Under Standard pricing, this equates to an effective price of approximately $0.10 per second.” The posted rates: video output tokens at $17.50 per million, text output at $9.00 per million, and input — text, image, video, or audio — at $1.50 per million. Run the vendor’s own math and it checks out: 5,792 tokens per second at $17.50 per million tokens is about $0.101 per second, which is the approximately-$0.10 figure Google states.
Here is the provenance detail that separates this report from most of the launch coverage: the four-tier per-second price table — $0.03 at 360p, $0.10 at 720p, $0.15 at 1080p, $0.30 at 4K — is not printed as a table on the Gemini API pricing page. The 720p effective rate is the only dollar-per-second figure Google states directly. The full table was reported, identically, by two separate outlets on launch day — XenoSpectrum and The Decoder — and it is internally consistent with Google’s published token rate. That makes the figures credible, but their correct label is “vendor-consistent, independently reported,” not “Google’s price sheet says.”
| Resolution | Google’s docs label (verbatim) | Reported price / second |
|---|---|---|
| Draft and default tiers | ||
| 360p — draft | “360p output resolution” | $0.03 |
| 720p — default | “720p output resolution (default)” | $0.10 |
| Upscaled tiers — per Google’s own parameter docs | ||
| 1080p | “1080p output (upscaled)” | $0.15 |
| 4K | “4k output (upscaled)” | $0.30 |
Sources for the table: resolution labels from Google’s Gemini API developer docs; per-second figures as reported by XenoSpectrum and The Decoder, August 27, 2026. The only vendor-stated dollar figure is the approximately-$0.10-per-second 720p rate, derived by Google from its published token rate. One price-context note from The Decoder’s comparison of Google’s own price sheet: the $0.30 per-second 4K tier matches what Google charges for Veo 3.1 Fast at 4K — same vendor, same rate, stated here without any quality comparison between the two.
Reported per-second output price by resolution tier
Source: XenoSpectrum + The Decoder, Aug 27, 2026 — reported figures, consistent with Google’s published token rateWhat we deliberately do not do here is convert these list rates into a cost per finished, client-ready second — that unit has to absorb draft passes, rejected takes, and extension chains, and it behaves very differently from a list price. We maintain a separate reference on normalizing video-model pricing to cost per finished second for exactly that exercise. For this launch, the list-price shape itself is the story: on the reported figures, a 10x spread separates drafting a second from delivering it at 4K — iteration and finishing priced as different activities.
06 — Platform StatusGA on one Google surface, Preview on another.
Here is the angle none of the launch-day coverage led with. On the direct Gemini Developer API, Google’s pricing page describes the model as “Our next-generation video generation and editing model, now generally available to developers on the paid tier of the Gemini API.” That is unambiguous GA language. Yet on the Gemini Enterprise Agent Platform — Google’s enterprise surface — the model page for the very same release lists a different model ID, gemini-omni-1.1-flash-preview, with the fields “Launch stage: Preview” and “Release date: August 27, 2026,” under Google’s boilerplate notice that Preview offerings are subject to its Pre-GA Offerings Terms.
Keep a second, easily-conflated fact separate from that split. Within the Developer API’s own naming, per XenoSpectrum, “Google lists gemini-omni-1.1-flash, used via the API, as the stable version, with the previous gemini-omni-flash-preview marked as preview” — that “-preview” suffix belongs to the older May-era model, not to 1.1. So there are two different preview labels in play on two different axes: the enterprise platform’s Preview stage for 1.1 itself, and the Developer API’s preview suffix for the superseded model. Conflating them produces confidently wrong statements about what is and is not production-ready.
The build-decision consequence is concrete. A team shipping against the direct Gemini API gets GA terms on day one. A team standardized on the enterprise platform is building on a Pre-GA offering with the contractual caveats that implies — even though it is the same underlying model from the same vendor on the same date. If your procurement or compliance process distinguishes GA from Preview, Omni 1.1 Flash is currently on both sides of that line depending on the door you enter through.
07 — Build DecisionsWhat changed, and where a team would feel it.
Assembled from Google’s launch post, the API docs, the enterprise platform page, and XenoSpectrum’s mechanism reporting — no single launch-day source put the before-and-after in one place, so we did.
| Capability | Omni Flash (May–Jun 2026) | Omni 1.1 Flash (Aug 27, 2026) | What it means for a build |
|---|---|---|---|
| Scene extension | Not in Omni — Google’s docs pointed developers to Veo 3.1 | 10-second passes to a 40-second cumulative ceiling, end-only | Longer sequences without leaving the Omni API; storyboard in 10-second beats |
| Extension context | Previous models referenced only the final second, per Google | Analyzes up to 10 seconds of prior footage per pass | Motion, lighting, and identity carry across the seam — chained clips stop looking chained |
| Keyframe control | Final-frame specification was a Veo 3.1 docs pointer | First-and-last-frame interpolation built in | Art-direct two stills; the model solves orbits, zooms, and loops between them |
| Draft workflow | No draft tier — new in 1.1 | 360p draft tier — a third of 720p cost, up to 60% faster by system throughput (vendor footnote) | Iterate at draft rates; reserve production tiers for approved cuts |
| Upscaled output | No upscale tier — new in 1.1 | 1080p and 4K available — both labeled “upscaled” in Google’s own docs | Delivery formats, not capture quality — validate upscaled output against your bar |
| Platform status | Now carries the “-preview” suffix on the Developer API | GA on the paid Gemini API; Preview on the enterprise platform under a separate model ID | Pick your surface deliberately — GA terms and Pre-GA terms coexist for the same model |
Where this leaves specific teams — with the caveat that this is a launch report, not a field ranking. For how the earlier Omni Flash sat against the rest of the video-model field before this release, see how the earlier Omni Flash compared against Sora and Veo and our July 2026 read across the video-model field.
Concept and pitch work
The 360p draft tier at a reported $0.03/s makes volume iteration defensible for the first time on this model line. Draft wide, review at 720p, upscale only the survivor.
Client-facing 4K work
Both top tiers are documented upscales. Run your own footage through the 1080p and 4K paths and judge against your delivery bar before writing 4K into a contract.
Upload-dependent workflows
Editing and extending externally uploaded videos faces additional regional restrictions, per XenoSpectrum’s read of Google’s terms. Verify current terms against your specific pipeline first.
Procurement-gated builds
The enterprise surface lists the model as Preview under Pre-GA terms while the direct API is GA. If compliance distinguishes the two, the direct Gemini API is the GA door today.
Looking forward: the draft-then-upscale split is the shape we expect video-model pricing to settle into across the industry, because it mirrors how production teams already work — cheap exploration, expensive finishing. Google recommending it as the workflow in its own Flow guidance can push the market toward pricing iteration and delivery as separate products. Teams that build their pipelines around that split now — rather than around one-shot generation at delivery resolution — are the ones positioned to benefit as the economics improve. If you are working out where AI video generation fits in your production or marketing stack, our AI transformation engagements and content engine work start with exactly this kind of workflow-and-cost mapping.
08 — ConclusionA workflow release wearing a headline number.
The real launch is the workflow, not the 40 seconds.
Omni 1.1 Flash’s headline reads like a capability jump — 40-second scenes, 4K output. Read the vendor’s own documents and it resolves into something more interesting: a workflow release. The 40 seconds is a chain of 10-second passes held together by a 10-second context window; the 4K is an upscale by Google’s own parameter docs; and the reported per-second tiers separate drafting and finishing at a 10x spread.
The honest reporting frame is the one Google’s own materials support: vendor-stated mechanics, a per-second price table that is independently reported and token-math consistent rather than printed on the price sheet, and no benchmark or quality claim of any kind in the launch materials themselves. Add the surface split — GA on the direct API, Preview on the enterprise platform — and the diligence for adopting this model is mostly about reading the right Google page for your situation.
The practical move is unchanged from every model launch we cover: run your own footage through the draft loop, judge the upscaled tiers against your own delivery bar, and price your pipeline on the billing mechanism — tokens — rather than the shorthand. Launch-day list prices are the start of that arithmetic, not the end of it.