AI product video crossed a pricing threshold in the first week of August 2026: the models that generate short product clips are now quoted per second, on published rate cards, in the same units you already use to budget API spend. That reframes catalog-scale video from a production question into an arithmetic one. A hundred SKUs at six seconds a clip is a two-figure line item — which means the hard part has moved somewhere else entirely.
Two of those rate cards landed inside a fortnight. Black Forest Labs took FLUX 3 Video to general availability on August 4, 2026 with vendor list pricing from $0.17 per second, a figure its OpenRouter listing matched the same day. MiniMax shipped H3 on July 31 with vendor pay-as-you-go pricing at $0.13 per second for 2K and $0.09 for 768P. Kling publishes a full per-second rate card too, but in yuan only. And ByteDance launched Seedance 2.5 on July 31 without publishing a per- second price at all.
This playbook does the thing none of the vendor pages do: it normalises those rates into one catalog-scale cost table, maps the pipeline from an existing product-detail-page image to a published ad, and gives equal billing to the failure modes — hands, on-screen text, and physics — that still make a human QA gate mandatory. If you have already worked through AI product photography for still images, this is the moving-picture sequel, and the economics are not the same shape.
- 01Per-second pricing is the change, not the models.FLUX 3 Video lists at $0.17/second as a vendor list price from its August 4, 2026 GA, matched by its OpenRouter listing; MiniMax H3 is $0.13/second at 2K and $0.09 at 768P on vendor pay-as-you-go. A 100-SKU run at six seconds a clip is $54.00 to $102.00 in generation cost.
- 02One headline model has no published price at all.Seedance 2.5 launched on July 31, 2026 with the launch post stating that ModelArk API access was coming soon. The per-minute figures circulating in coverage are third-party estimates, not ByteDance-quoted — treat them as unpriced until the vendor says otherwise.
- 03Kling’s rate card is denominated in yuan only.Kling 3.0 runs ¥0.6 to ¥3.0 per second depending on resolution and whether audio is included. No USD equivalent is published, so any dollar figure — including ours — is an FX conversion you did, not a vendor list price.
- 04The QA gate costs more than the generation.On the stated assumptions below — three minutes of human review per clip, an illustrative $40/hour blended internal rate — review labour on a 100-SKU run runs roughly double the FLUX 3 Video generation bill and close to four times the cheapest MiniMax tier.
- 05Hands, on-screen text and physics still break.No vendor publishes a product-video accuracy benchmark, and none of the three claims these are solved. Reflective surfaces, liquid pours, fabric drape and small label text are the recurring failure points — a pattern across launch materials, not a scored evaluation.
01 — What ChangedThe unit of cost is now a second, not a shoot day.
For as long as product video has existed, its unit of cost has been a day — a studio day, a crew day, an edit day. That unit is indivisible and it does not scale down, which is why catalog-scale video has historically been reserved for hero SKUs and seasonal campaigns. Everything else got a still image and a bullet list.
Three of the four models below broke that unit into seconds — the fourth is the counterexample this post keeps coming back to. Not because the output is equivalent to a studio shoot — it is not, and section 05 is entirely about where it falls short — but because a divisible unit lets you ask a different question. Instead of “which twelve SKUs deserve video,” the question becomes “what does it cost to give all 400 of them a six-second clip, and what does the review process cost on top.”
FLUX 3 Video
Black Forest Labs’ first general availability for video. Clips up to 20 seconds, HD 720p native with Full HD 1080p via upscaling, image-to-video with multiple keyframes, native audio and dialogue with lip-sync in around 14 languages, plus a Draft Mode that previews cheaply before committing to a full render.
MiniMax H3
A video-generation model, not a language model — and the model ID is MiniMax-H3, never the community label some coverage uses. Reference-to-video from up to nine images, three videos and three audio clips, native synced audio, and instruction-based editing of an existing clip. Prepaid Hailuo subscription packages explicitly exclude H3.
Kling 3.0
The most granular rate card of the four: separate per-second rates for Turbo with audio, base silent, and Omni, tiered by resolution, with 4K published on the silent tiers. One credit equals one yuan. The page publishes no dollar equivalent, so a USD figure is always a conversion someone else applied.
Seedance 2.5
The deepest reference ceiling in the field — up to 30 images, 10 videos and 10 audio clips, 50 references in total — with 30-second single-pass generation extendable across rounds into multi-minute output. The launch post stated that ModelArk API access was coming soon, and no per-second rate accompanied it.
Worth separating two events that coverage keeps merging: the FLUX 3 announcement in late July was early access to the base FLUX 3 image and multimodal family, and it is a different thing from the August 4 video GA. The 20-second clips, the native audio, and Draft Mode all belong to the video release. Our FLUX 3 Video GA guide covers that launch in depth, and the MiniMax H3 launch and Seedance 2.5 one-take launch posts do the same for the other two.
02 — Proprietary AnalysisThe rate cards, normalised to a 100-SKU run.
Every vendor quotes its own number, in its own currency, on its own surface. Nobody publishes the comparison, so we built it. The table below takes a fixed workload — 100 SKUs, one six-second clip each — and applies each published rate to it. Every cell is arithmetic on a rate we could read on a primary page at the time of writing: price per second × 6 for the clip, × 100 for the run. No efficiency claims, no discounts, no volume tiers folded in.
Read the surface label on every row, because they are not interchangeable. A vendor list price is what the model owner charges directly; an aggregator listing is what a routing layer charges to reach the same model. On FLUX 3 Video the two agree on the base rate quoted here — but that is something to check per model and per tier rather than assume, and either surface can change without an announcement.
| Model · tier | Rate per second | One 6s clip | 100-SKU run | Native audio | Max single pass |
|---|---|---|---|---|---|
| USD-denominated vendor rates · list price or pay-as-you-go, read from the vendor page itself | |||||
| FLUX 3 Video · 720p native | $0.17 · vendor list | $1.02 | $102.00 | Yes · dialogue + lip-sync | 20s |
| MiniMax H3 · 2K | $0.13 · vendor PAYG | $0.78 | $78.00 | Yes · synced | 15s |
| MiniMax H3 · 768P | $0.09 · vendor PAYG | $0.54 | $54.00 | Yes · synced | 15s |
| CNY-denominated vendor rate card · dollar figures are our illustrative ~7:1 conversion, never a vendor list price | |||||
| Kling 3.0 Turbo · 1080P, audio | ¥1.0 | ¥6.00 (~$0.86) | ¥600 (~$86) | Yes | Not on the rate card |
| Kling 3.0 Turbo · 720P, audio | ¥0.8 | ¥4.80 (~$0.69) | ¥480 (~$69) | Yes | Not on the rate card |
| Kling 3.0 · 720P, silent | ¥0.6 | ¥3.60 (~$0.51) | ¥360 (~$51) | No | Not on the rate card |
| Kling 3.0 · 1080P, silent | ¥0.8 | ¥4.80 (~$0.69) | ¥480 (~$69) | No | Not on the rate card |
| Kling 3.0 · 4K, silent | ¥3.0 | ¥18.00 (~$2.57) | ¥1,800 (~$257) | No | Not on the rate card |
| No vendor price published | |||||
| Seedance 2.5 | Not published · API access stated as coming soon at launch | — | — | Yes | 30s, extendable |
Two things jump out of that grid. First, the USD spread across the comparable tiers is narrow: $54.00 to $102.00 for the same 100-SKU, six-second workload. At that spread, model choice should be driven by output quality and reference handling, not by price — a $48 difference across an entire catalog run is noise next to one round of reshoots. Second, the 4K silent tier on Kling is roughly five times the 720P silent tier per second, which is the only place in this table where a resolution decision materially moves the budget.
Generation cost · 100 SKUs at one 6-second clip · USD-denominated rates only
Our calculation: rate per second × 6 seconds × 100 SKUs, from rates published at the time of writingOne more surface note, because it changes how you shop. OpenRouter is not a universal price-comparison layer for video the way it is for text models — at the time of writing its listing carried FLUX 3 Video but not Kling 3.0, MiniMax H3 or Seedance 2.5, each of which is API-direct or vendor-platform only. If you are used to comparing language models on one aggregator page, expect to visit three or four vendor consoles instead. The broader dynamics of that market are covered in our piece on the AI video price war.
03 — The PipelineFrom a product-page image to a published clip.
The pipeline that actually works at catalog scale does not start from a text prompt. It starts from the photography you already have on the product detail page, because that is the only asset that guarantees the generated clip shows your product rather than a plausible cousin of it. Every model in this comparison supports image-to-video or reference-to-video for exactly this reason.
Below is the stage map, with the model feature that serves each stage and the point where that stage typically breaks. It is not a vendor diagram; it is the sequence a team ends up running once the first batch comes back with problems.
| Stage | What happens | Model feature that applies | Typical failure point |
|---|---|---|---|
| PDP image → short clip → distribution surface | |||
| 01 · Source image prep | Pick the cleanest existing PDP frame per SKU. Normalise crop, background and colour so the batch is consistent before anything is generated. | Image-to-video entry point on every model in the table | Inconsistent source lighting propagates into every clip in the run |
| 02 · Reference selection | Choose the keyframes or reference set that pins the product’s identity — shape, finish, logo placement — across the shot. | FLUX 3 Video multi-keyframe I2V · MiniMax H3 reference-to- video · Seedance 2.5 multi-reference inputs | Too few references and the model drifts; too many conflicting ones and it averages them |
| 03 · Generation | Render a cheap preview first, judge subject and motion, then commit the full render at the resolution the surface needs. | FLUX 3 Video Draft Mode, which preserves subject, composition and motion from draft to final render | Committing full renders before the motion is right — the single most expensive habit at volume |
| 04 · QA gate | Human review against a fixed checklist: hands, on-screen text and labels, and physically plausible motion for liquids, fabric and reflective surfaces. | None — no vendor publishes a product-video accuracy benchmark to automate against | Skipping the gate on “easy” SKUs, which is exactly where label text quietly garbles |
| 05 · Format conversion | Re-crop and re-encode per destination: PDP embed, Meta Feed at 4:5, and vertical short-form placements. | Post-process, not a generation setting — crop and re-encode from one render per SKU | Regenerating per aspect ratio instead of cropping, which multiplies both cost and QA load |
04 — SKU ConsistencyKeeping one SKU looking like itself.
Consistency is the whole ballgame in ecommerce video. A model that produces a beautiful clip of something that is 90% your product has produced nothing you can publish. The lever the three priced models give you is the reference system: how many anchoring inputs you can supply, and of what kind. They differ enough that this alone can decide the model.
Keyframes plus Draft Mode
Image-to-video with multiple keyframes, which is the direct fit for a pipeline that starts from existing PDP photography rather than a prompt. Draft Mode renders a fast, cheap preview and then a full render that keeps the same subject, composition and motion — a cost-control mechanic built for iterating at volume. Native audio and dialogue with lip-sync in around 14 languages make localised spokesperson clips viable from the same pipeline.
Reference-to-video, plus editing
Reference-to-video generation from up to nine images, three videos and three audio clips, with native synced audio and instruction-based editing applied to an existing clip. That editing path matters at catalog scale: a near-miss clip can sometimes be corrected rather than regenerated, which is a different unit economics from a full re-render.
The deepest reference ceiling
Up to 30 images, 10 videos and 10 audio clips — 50 references in total, the widest anchoring surface of the models compared here — with 30-second single-pass generation extendable across multiple rounds into multi-minute output. The catch is commercial rather than technical: no per-second price has been published, so you cannot yet put it in a budget.
The practical read: if your catalog has strong existing photography and each SKU needs one clean clip, keyframe-driven image-to-video is enough and FLUX 3 Video’s Draft Mode is the cost lever. If your SKUs are visually similar to each other — variants of the same silhouette in different finishes — you want the widest reference set you can get, because that is what stops the model averaging across your own range. And if the plan is one master creative with the product swapped per variant, that is a different technique entirely, covered in our guide to swapping product creative for local ad variants.
05 — What Still FailsHands, on-screen text, and physics.
Here is the honest framing, and it needs stating precisely because it is easy to overclaim in both directions. There is no published benchmark for product-video accuracy. None of the vendors scores hands, small on-screen text, or physical plausibility; none claims those problems are solved. What exists is a pattern visible across the vendors’ own launch materials and early independent coverage — the same categories of artefact keep appearing, across models.
Treat the list below as a QA checklist, not as a measured error rate. The distinction matters when you are setting expectations with a merchandising team: “these are the things to look for” is defensible; “model X fails 12% of the time on reflective surfaces” is a number nobody has published.
The three recurring categories
- Hands. Any clip where a person holds, opens, or demonstrates the product puts hands in frame at close range — the single most scrutinised region in generated video, and the one a customer will notice first.
- On-screen text. Labels, ingredient panels, model numbers, and brand marks on packaging. This is the failure mode with the highest commercial consequence, because a garbled ingredient list on a consumable is not an aesthetic problem.
- Physics. Liquid pours, fabric drape, and reflective or transparent surfaces. Glass, chrome, foil packaging and anything with a specular highlight are where implausible motion is most visible.
“The generation bill is the cheap part. The QA gate is the line item nobody puts on the rate card.”— Digital Applied, on catalog-scale video pipelines
06 — The Real CostThe QA gate is the line item that is not on the rate card.
Section 02 priced generation. This section prices the part that decides whether the project is worth running. The arithmetic below is planning arithmetic on stated assumptions — three minutes of human review per clip, an illustrative $40/hour blended internal rate, and a quarter of clips failing the gate and needing one regeneration. Those assumptions are yours to argue with; substitute your own numbers and the shape of the answer holds.
One 6-second clip each
100 SKUs × 6 seconds × $0.17 per second, the FLUX 3 Video vendor list rate at the time of writing. On MiniMax H3’s 768P tier at $0.09 per second the identical run is $54.00. This is the number every vendor page lets you compute — and the least important one in the model.
Three minutes a clip
100 clips at three minutes of human review each is 300 minutes, or five hours. At an illustrative $40/hour blended internal rate that is $200 — roughly double the FLUX 3 Video generation bill and close to four times the cheapest MiniMax tier. Nobody publishes this number because nobody sells it to you.
The number to actually plan against
Assume a quarter of clips fail the gate and need one regeneration: 25 extra clips add $25.50, taking generation to $127.50. Re-reviewing those 25 adds 75 minutes, taking review to 6.25 hours, or $250. Total $377.50 — with generation at roughly a third of it, and human time at the rest.
That ratio is the finding, and it inverts the intuition most teams bring to the decision. The instinct is to shop hard on per-second price; the arithmetic says the per-second price is roughly a third of the true cost and the spread between the cheapest and most expensive comparable model is $48 across an entire 100-SKU run. The variable that actually moves your budget is the failure rate, and the variable that moves the failure rate is the quality and consistency of your source images — stage 01, the one that costs nothing but attention.
Which suggests a different procurement question. Not “which model is cheapest per second,” but “which model’s reference system most reduces the number of clips my reviewers reject.” A model that costs $0.08 more per second and halves your rejection rate is straightforwardly the cheaper option on this arithmetic. That is not a question any rate card answers, and it is the one worth running a paid pilot on — twenty SKUs across two models, one reviewer, same checklist. It is the shape of evaluation our ecommerce engagements start with.
07 — DistributionWhere the clips actually run.
A generated clip is not a deliverable until it fits the surface it runs on, and the surfaces disagree with each other — and, in one case, with the models. Shopify’s own product-video guidance puts the sweet spot at 30 to 60 seconds, with 15 seconds to 2 minutes as the usable range. FLUX 3 Video caps a single pass at 20 seconds and MiniMax H3 at 15. That gap is not a flaw in either; it just means a 30-second product-page video is a stitched sequence or a multi-round extension, not one generation.
The product page itself
Shopify’s guidance puts the product-video sweet spot at 30 to 60 seconds, with 15 seconds to 2 minutes usable. That is above the single-pass ceiling of FLUX 3 Video (20s) and MiniMax H3 (15s), so plan a 30-second PDP clip as a stitch of shorter passes or a multi-round extension — and treat longer explainer and tutorial content as a separate production entirely.
Meta Feed, 4:5
Meta’s Feed video spec accepts MP4, MOV or GIF, durations from 1 second to 241 minutes, files up to 4GB, H.264 encoding with square pixels and stereo AAC audio at 128kbps or better. The recommended Feed resolution is 1440×1800 at a 4:5 ratio, with a 120×120 minimum.
Vertical placements
Vertical feeds want a different crop from the 4:5 Feed asset, which makes aspect ratio a rendering decision at the end of the pipeline rather than a generation setting. Confirm the current vertical spec on the platform’s own ads guide before you batch-render — those numbers change, and a stale spec sheet is an expensive way to find out.
Draft before you commit
FLUX 3 Video’s Draft Mode renders a cheap preview and then a full render that preserves the same subject, composition and motion. On a single clip that is a convenience; across 100 clips it is the difference between paying once and paying three times for the same shot. Build the draft-then-commit loop into the pipeline before the first batch, not after it.
One clip therefore serves several surfaces if — and only if — you re-crop it in post rather than re-render it per placement. Regenerating per aspect ratio multiplies both the generation bill and, more painfully, the QA load, because every re-render is a new clip that has to clear the gate. For the short-form end of the funnel, our guide to short-form product clips on YouTube Shorts covers the platform-specific side, and if 4K masters are the goal, Kling 3’s 4K rendering is the tier to price against. On the buying side, the placement mechanics sit with paid media, not with the generation stack.
08 — The Business CaseDoes it pay back?
The demand-side evidence for product video is widely quoted and weakly sourced, and it is worth being precise about what it is. A 2026 Wyzowl marketing survey reports that 85% of people say they have been convinced to buy a product or service by watching a video, with 63% preferring a short video to learn about a product against 12% for text articles. A Shopify guide puts the equivalent purchase-influence figure at 89%, and reports that 96% have watched an explainer video to learn about a product or service.
Both are marketing content produced by a video-marketing vendor and an ecommerce platform respectively — not independently audited research, and not academic work. Nor can we show that the two are independent of each other: the figures land close together, but a platform guide quoting a purchase-influence number may be restating the same vendor survey, which would make the agreement one origin counted twice rather than two readings that match. Treat it as a single directional signal. What it is not useful for is a conversion forecast for your catalog. Use it to justify a pilot, never to size a business case.
The defensible business case is the one you can compute from the numbers in section 06. A 100-SKU run costs on the order of $377.50 all-in on the stated assumptions. Against that, the only question that matters is whether the clips move the metric on the pages they sit on — which is a measurable thing, on your own traffic, with a holdout set of SKUs that stay video-free. That test costs the same $377.50 and produces a number that belongs to you rather than to a vendor’s survey panel. Wire it into the PDP itself as a controlled experiment and the result generalises to the rest of the catalog.
Where this goes next
Two pressures are visible in the current field and both point the same way. First, per-second rates across the comparable tiers have converged into a narrow band, which is what happens when a capability commoditises; the differentiation is migrating to reference handling, editing of existing clips, and cost-control mechanics like draft renders. Second, the single-pass ceiling — 15 to 30 seconds today — sits just below the length the distribution surfaces actually want, which is an obvious gap for the next release cycle to close.
If both hold, the bottleneck in catalog-scale product video will not be model access or price by the end of 2026. It will be review throughput: the human hours between a rendered clip and a published one. Teams that invest now in a tight, checklist-driven QA gate and clean, normalised source photography will be able to absorb the next price drop as pure volume. Teams that treat generation cost as the project will find that the cheaper the seconds get, the more the review queue becomes the entire problem.
09 — ConclusionThe seconds got cheap. The judgement did not.
Generation is a third of the cost. Review is the rest — and the only part you control.
The pricing story is real and it is simple. FLUX 3 Video reached general availability on August 4, 2026 at $0.17 per second on its vendor list; MiniMax H3 shipped four days earlier at $0.09 to $0.13 on vendor pay-as-you-go; Kling publishes a granular per-second card in yuan. A hundred SKUs, six seconds each, is $54 to $102 of generation. That is not a budget line anyone needs to defend.
The part worth defending is everything around it. Seedance 2.5 has no published per-second price, so the figures circulating in coverage are third-party estimates rather than vendor numbers. Kling has no published dollar price, so every USD figure — including the approximations in our own table — is a conversion, not a quote. And no model publishes an accuracy benchmark for the three things that actually break a product clip: hands, on-screen text, and physics.
So the operating posture is straightforward. Normalise your source photography before you generate anything. Draft before you commit to a full render. Put a fixed checklist between the render and the publish button, and staff it honestly — on the arithmetic here, it is two-thirds of the true cost. Then run the only test that generalises: a holdout set of SKUs on your own product pages, measured on your own traffic. The vendor surveys are directionally encouraging. Your own holdout is the only number you can act on.