MarketingIndustry Guide12 min readPublished July 24, 2026

Two video lines, one company · audio-native flagship vs open workhorse · from $0.10/sec depending on host

Inside Alibaba's Video Stack: HappyHorse 1.1 + Wan 2.7

Alibaba's video effort is two products, not one. HappyHorse is the audio-native flagship that debuted anonymously at the top of the blind-test arena in April and shipped its 1.1 upgrade on June 21–22. Wan 2.7 is the cheaper, more editable, partially open workhorse from Tongyi Lab, live since April 6. Both are established lines by now — this is a stack review for creative-production teams, not launch news.

DA
Digital Applied Team
Senior strategists · Published Jul 24, 2026
PublishedJul 24, 2026
Read time12 min
SourcesCNBC, VentureBeat, TechNode, fal.ai
Arena Elo · Jun 22 snapshot
1,444
T2V & I2V · No. 2 (Arena.ai)
Jul 17 read: 1,149 T2V
Wan 2.7 on fal.ai
$0.10/sec
flat, pay-per-use
Wan 2.7 architecture
27B
total · 14B active MoE
Apache 2.0
HappyHorse lip-sync
7langs
joint audio-video · carried-over 1.0 spec

Alibaba's AI video stack is easy to misread as one product, but it is two distinct model lines run in parallel: HappyHorse, the audio-native flagship that appeared anonymously at the top of the Artificial Analysis Video Arena in April 2026, and Wan 2.7, the cheaper, more editable suite from Tongyi Lab with Apache 2.0 open weights for its base models. Neither is new — HappyHorse 1.1 shipped June 21–22 and the Wan 2.7 line launched April 6 — which is exactly why a stack-level read is now more useful than another launch recap.

The stakes for creative-production teams are practical, not academic. The two lines price differently on the same host, expose different generation modes, and sit at opposite ends of Alibaba's openness spectrum. Meanwhile the competitive field around them has contracted — VentureBeat's June analysis framed OpenAI's discontinued standalone Sora product and ByteDance's shelved international Seedance rollout as the backdrop against which Alibaba's stack now competes.

This guide covers how HappyHorse went from anonymous arena entry to claimed flagship, what 1.1 actually ships, how Wan 2.7 differs and why it costs less, why the two published arena Elo snapshots do not reconcile, how pricing varies by host, and — the part nobody else is writing — why the same company ships three different openness postures across three modalities.

Key takeaways
  1. 01
    Two parallel video lines, neither of them new.HappyHorse 1.1 shipped June 21–22, 2026; the Wan 2.7 suite launched April 6 with incremental snapshots through June 12. Neither is launch news as of late July — treat this as a stack decision, not a release chase.
  2. 02
    HappyHorse debuted anonymously — and won.An unbranded model took the No. 1 spot on both text-to-video and image-to-video blind-test leaderboards around April 7, 2026. Alibaba confirmed authorship three days later, on April 10, per CNBC.
  3. 03
    Arena standing has drifted, and the snapshots don't reconcile.VentureBeat, citing Arena.ai on June 22, reported HappyHorse at Elo 1,444 and No. 2 across leaderboards. A July 17 Hedra read lists 1,149 T2V / 1,111 I2V. Different platforms, dates, and vote pools — never merge them into one figure.
  4. 04
    Pricing varies by host — roughly $0.10 to $0.28 per second.On fal.ai, Wan 2.7 runs $0.10/sec flat while HappyHorse runs $0.14/sec at 720p and $0.28/sec at 1080p. Hedra's July comparison lists different per-minute rates. There is no single canonical price.
  5. 05
    Alibaba's openness posture is inconsistent by design or by accident.Wan ships Apache 2.0 weights and published architecture. HappyHorse ships neither weights nor a tech report. Qwen-Image-3.0 went invite-only with no benchmarks. Three postures, one company, one April-to-July window.

01The StackOne company, two video lines.

Early speculation on X held that HappyHorse might secretly be Wan 2.7 under a codename. It is not — they are separate models from separate teams, and comparison leaderboards list them as separate entries. HappyHorse is the closed flagship; Wan 2.7 is Tongyi Lab's suite, with open weights for its base models and a published architecture. Understanding them as a two-tier stack — audio-native hero model on top, cheaper editable workhorse below — is the single most useful framing for a production team.

The field they compete in has thinned. For the wider post-Sora landscape and how Google's and xAI's entries stack up, see our AI video field comparison; for the July snapshot of FLUX 3, Seedance 2.5, and Gemini's video line, the FLUX 3-led video field comparison published alongside this post covers the non-Alibaba half of the market.

Flagship line
HappyHorse 1.1
Audio-native · closed · June 21–22, 2026

Joint audio-video generation in a single pass, lip-sync across seven languages, four generation modes. No open weights, no official tech report — the 15B unified-Transformer architecture is community-documented.

HappyHorse site · Bailian · Model Studio API
Workhorse line
Wan 2.7
Open-leaning · Tongyi Lab · April 6, 2026

27B total / 14B active MoE, Apache 2.0 license for base models, four sub-models and seven generation modes including natural-language video editing. Cheapest per second of the two on the same host.

Apache 2.0 · fal.ai · Model Studio
Market backdrop · VentureBeat, June 22, 2026
VentureBeat's June analysis opened with a stark framing of the field: "The AI video generation market entered 2026 with three credible enterprise contenders. One is dead. One is frozen. And the one still standing is a Chinese company backed by $52.7 billion in infrastructure spending…" — referencing OpenAI's discontinued standalone Sora product and ByteDance's indefinitely shelved international Seedance 2.0 rollout. Both are backdrop, not news — but they explain why Alibaba's two-line stack now draws so much enterprise attention.

02Origin StoryThe anonymous arena debut.

HappyHorse's origin is the most unusual go-to-market in the 2026 video-model field. Around April 7, 2026, an unbranded model calling itself "HappyHorse" appeared on the Artificial Analysis Video Arena — the blind-test platform where voters compare clips without knowing which model produced them — and immediately took the No. 1 spot on both the text-to-video and image-to-video leaderboards, ahead of rivals including ByteDance's Seedance line. Three days later, on April 10, CNBC reported Alibaba's confirmation that the model was its own.

Attribution below the company level is thinner. Trade press reports place the model with an Alibaba Token Hub (ATH) team with roots in the Taobao/Tmall Future Life Lab, restructured into a standalone unit — though outlets word the org chart differently, and Alibaba has not published its own account. The same trade coverage names Zhang Di, reported as a roughly 15-year AI veteran, former Kuaishou VP, and technical architect of Kling AI who rejoined Alibaba in late 2025, as the team's lead. Neither the exact lab name nor the team-lead attribution appears in CNBC's or VentureBeat's primary reporting, so treat both as trade-press-reported rather than Alibaba-confirmed.

The anonymous-debut strategy matters beyond trivia. Winning a blind test before revealing the brand neutralized the discount that Western evaluators sometimes apply to Chinese-lab releases — voters ranked the output first and learned the vendor afterward. It is also, notably, the inverse of how Alibaba launched Qwen-Image-3.0 three months later: maximum benchmark transparency first versus no benchmarks at all. That inconsistency is Section 07's subject.

Arena debut
Anonymous No. 1
Apr 7

An unbranded 'HappyHorse' entry took the top spot on both text-to-video and image-to-video blind-test leaderboards on the Artificial Analysis Video Arena, per CNBC's reporting.

Both leaderboards
Authorship claimed
Alibaba confirms
+3days

On April 10, 2026, Alibaba confirmed it built HappyHorse — three days after the anonymous debut. The blind-test-first strategy meant voters ranked output before knowing the vendor.

CNBC, Apr 10, 2026
First API partner
fal.ai goes live
Apr 27

fal launched developer and enterprise access to HappyHorse-1.0 as the first official third-party API partner, with four endpoints: text-to-video, image-to-video, reference-to-video, and video-edit.

4 endpoints

03The FlagshipWhat HappyHorse 1.1 actually ships.

HappyHorse 1.1 arrived June 21–22, 2026 — Alibaba Cloud's own post landed Sunday June 21, with TechNode's write-up following on June 23. Alibaba billed it as a major upgrade over 1.0, with vendor-stated improvements in motion dynamics, subject consistency, prompt adherence, visual quality, and audio generation. No independent before/after benchmark has been published, so those deltas remain the vendor's own claims.

The mode lineup grew to four: text-to-video, image-to-video, subject-to-video, and — new in 1.1 — video editing, all with integrated audio at no additional cost through a single API endpoint. The carried-over 1.0 spec is the differentiator: 1080p output, 3–15 second clips, five aspect ratios, and native lip-sync across seven languages (Mandarin, Cantonese, English, Japanese, Korean, German, French). Architecturally, HappyHorse generates audio and video jointly in a single pass through a unified 40-layer self-attention Transformer with no cross-attention modules — per fal's product documentation — rather than stitching a video model to a separate audio model. Secondary trade coverage additionally describes synchronized dialogue, ambient sound, and effects in one inference step at a reported 24fps, though those specifics do not appear verbatim on Alibaba's or fal's own pages.

The parameter figure deserves its own hedge: the widely cited 15-billion-parameter unified self-attention Transformer — one token sequence carrying text, image, video, and audio tokens — comes from community-compiled technical documentation referenced by VentureBeat, because Alibaba has published no official HappyHorse paper. Availability is broad regardless: the HappyHorse website, Alibaba Cloud's Bailian platform, and Qwen Cloud for end users, with full API access via Alibaba Cloud Model Studio for enterprise customers. A 40% sitewide discount ran for the first two weeks after the 1.1 release — expired by this writing.

Generation modes
Video editing is new
4

Text-to-video, image-to-video, subject-to-video, and video editing — the last added in 1.1. All four carry integrated audio at no extra cost via one API endpoint.

One endpoint
Lip-sync languages
Joint audio-video
7

Mandarin, Cantonese, English, Japanese, Korean, German, French — generated in the same single pass as the video rather than dubbed on afterward.

Single-pass
Clip envelope
1080p, five ratios
3–15sec

Output up to 1080p in 16:9, 9:16, 1:1, 4:3, and 3:4 — the short-form envelope that covers most paid-social and organic placements without cropping.

Per fal.ai spec

04The WorkhorseWan 2.7: the open workhorse.

Wan 2.7 is the older, better-documented, and cheaper half of the stack. Tongyi Lab launched the full suite on April 6, 2026 — the image component surfaced around April 1 and text-to-video was live by April 3 — with incremental snapshot updates continuing through June 12. Where HappyHorse ships as a sealed product, Wan ships as an architecture: 27 billion total parameters with 14 billion active per step via mixture-of-experts, under an Apache 2.0 license, split across four sub-models — text-to-video, image-to-video, reference-to-video, and video edit.

The capability envelope trades audio-native polish for breadth. Clips run 2–15 seconds at 720p or 1080p across seven generation modes, including start–end frame animation, video continuation, audio-to-video, and multi-reference consistency. The most production-relevant feature is natural-language editing of nearly every video element — character actions, dialogue, appearance, scene, style, camera — which turns revision rounds from re-rolls into edits. For iteration-heavy marketing work, that editing surface plus the lower per-second price is often worth more than HappyHorse's superior sync.

Why the license matters
Apache 2.0 on Wan's base models means a team can self-host, fine-tune on brand assets, and build products on top without per-second API economics — the same calculus that made open-weight language models an enterprise category. HappyHorse offers none of that: no weights, no tech report, and an architecture known only through community documentation. Same company, opposite postures.

05BenchmarksTwo Elo snapshots that don't reconcile.

Most coverage repeats one of two stories in isolation: the April "anonymous No. 1" story or the June "No. 2" story. Put the published numbers side by side, though, and an honest caveat emerges — the two most-cited Elo snapshots cannot be stitched into a single trend line, and anyone using "arena rank" in a buying decision should understand why.

Jun 22, 2026 snapshot
No. 2 across leaderboards
1,444

VentureBeat, citing Arena.ai, reported HappyHorse at 1,444 Elo in both text-to-video and image-to-video — leading Google's Veo-3.1 (with audio) by 69 points in T2V and xAI's Grok-Imagine-Video by 23 points in I2V.

Arena.ai via VentureBeat
Jul 17, 2026 snapshot
HappyHorse 1.1 · T2V
1,149

Hedra's model-comparison post lists HappyHorse 1.1 at 1,149 T2V / 1,111 I2V on 3,535 votes, calling it a serious #5 alternative at #4 overall and #4 on text-to-video. A different platform read, weeks later — not the June number minus decline.

Hedra blog · 3,535 votes
Jul 17, 2026 snapshot
Wan 2.7 · No. 3 T2V
1,159

The same Hedra read places Wan 2.7 at 1,159 Elo, No. 3 on the text-to-video leaderboard on 3,153 votes — the cheaper sibling out-ranking the flagship on that particular snapshot.

Hedra blog · 3,153 votes

The 1,444 and 1,149 figures are not directly comparable. Arena Elo drifts as new models enter and vote pools accumulate, and the two sources may be reading different leaderboard categories at different moments. What the pair of dated snapshots supports is a qualitative claim: HappyHorse's standing became more contested between late June and mid-July as Veo 3.1 and other rivals closed the gap, and by Hedra's July read Wan 2.7 ranked above its own flagship sibling on text-to-video.

The interpretive lesson generalizes past Alibaba: arena rank is a moving, provider-dependent measurement, not a fixed property of a model. A leaderboard position quoted without a date and source is close to meaningless for procurement. If a video model's rank is load-bearing in your vendor decision, capture the snapshot — platform, date, category, vote count — the way you would any other perishable benchmark, and re-check it before renewal.

06EconomicsPricing varies by host, not just by model.

There is no single canonical price for either model — rates are set independently by each hosting provider. On fal.ai, HappyHorse's first official third-party API partner, HappyHorse runs $0.14 per second at 720p and $0.28 per second at 1080p, while Wan 2.7 runs a flat $0.10 per second at either resolution — a ten-second 1080p clip costs $1.00 on Wan versus $2.80 on HappyHorse, a 2.8× gap that narrows to 1.4× at 720p. Hedra's July 17 comparison lists different rates again: $9.00 per minute for Wan 2.7 at 1080p (≈$0.15/sec) and $9.90 per minute for HappyHorse 1.1 (≈$0.165/sec). All rates are pay-per-use with no minimums on fal.

Per-second cost · Alibaba video models across hosts

Sources: fal.ai product pages (retrieved Jul 24, 2026); Hedra model-comparison blog (Jul 17, 2026). Rates set independently per host.
Wan 2.7 · fal.ai$0.10/sec flat · 720p or 1080p
$0.10
HappyHorse · fal.ai · 720p$0.14/sec · pay-per-use
$0.14
Wan 2.7 · Hedra-listed · 1080p$9.00/min ≈ $0.15/sec · Jul 17 snapshot
$0.15
HappyHorse 1.1 · Hedra-listed · 1080p$9.90/min ≈ $0.165/sec · Jul 17 snapshot
$0.165
HappyHorse · fal.ai · 1080p$0.28/sec · pay-per-use
$0.28

Two readings follow. First, the spread across hosts — HappyHorse 1080p at $0.28/sec on fal versus roughly $0.165/sec in Hedra's listing — means there is no single canonical rate for the same model, which argues for re-quoting rates at the start of every production cycle rather than baking one listing's number into a budget. (Hedra publishes a comparison; its role as a host is not confirmed, so treat its figures as a listing, not a competing provider's price.) Second, the model gap on a single host is large enough to shape workflow: at fal's rates, every hero-quality HappyHorse second at 1080p buys 2.8 seconds of Wan output. Teams that iterate — which is every team — can afford nearly three times the draft volume on the workhorse before committing the flagship to final renders.

07The PatternThree modalities, three openness postures.

Here is the angle the launch coverage misses: between April and July 2026, the same company shipped three flagship-tier generative lines with three incompatible openness postures. Wan 2.7 shipped with Apache 2.0 weights for its base models and a published architecture. HappyHorse shipped with no weights and no tech report — its architecture exists publicly only as community-compiled documentation. And Qwen-Image-3.0, launched July 21, went further still: API and invite-only, with no weights, no benchmarks, no license, and no tech report, a sharp break from Qwen-Image 1.0's Apache 2.0 release with a published report in August 2025.

Video · Wan 2.7
Open posture
Apache 2.0 · architecture published

27B/14B MoE detailed publicly, base-model weights downloadable, four sub-models documented. The posture that built Alibaba's open-model credibility.

Tongyi Lab · Apr 6, 2026
Video · HappyHorse
Sealed posture
No weights · no tech report

Benchmark-transparent (it debuted on a public blind-test arena) but artifact-closed: the 15B unified-Transformer description exists only as community documentation.

ATH team (trade-press-reported)
Image · Qwen-Image-3.0
Closed posture
Invite-only · no benchmarks · no license

Launched Jul 21, 2026 with striking capabilities — 4,500-token prompts, legible 10px text — and zero published evidence, reversing Qwen-Image 1.0's open release.

Sharp break from Aug 2025

We covered the image-side retreat in detail in Alibaba's broader closed-flagship pivot and Qwen-Image-3.0's own closed pivot. Read against the video stack, the pattern looks less like a single strategy shift and more like per-team autonomy: the Wan and Qwen open-weight tradition on one side, and newer flagship-chasing units on the other, each choosing its own disclosure level. For buyers, the practical consequence is that "Alibaba is an open-model company" is no longer a safe planning assumption — openness now has to be verified per product line, per release.

Looking forward, two signals are worth watching. First, whether third-party hosts that carried HappyHorse 1.0 update to 1.1 — VentureBeat explicitly framed that as an open indicator of developer demand as of June 22. Second, whether HappyHorse ever gets a tech report or weights: if it does, the sealed posture was a competitive-window tactic; if it doesn't, Alibaba's video stack settles permanently into a closed-flagship-plus-open-workhorse shape — which, notably, is the same two-tier structure most Western labs have converged on.

08Decision MatrixWhich model for which job.

No swept source puts both hosts' pricing and both models' specs side by side — coverage treats HappyHorse origin stories and Wan spec sheets as separate beats. The table below is that missing head-to-head, compiled from fal.ai's product pages, Hedra's July comparison, TechNode, and VentureBeat.

Head-to-head comparison of Alibaba's two video model lines, HappyHorse 1.1 and Wan 2.7, across audio capability, clip envelope, architecture, openness, generation modes, per-second pricing on fal.ai and Hedra-listed rates, and best-fit use case. Compiled from fal.ai product pages and Hedra's model comparison (July 17, 2026), TechNode (June 23, 2026), and VentureBeat (June 22, 2026), retrieved July 24, 2026.
SpecHappyHorse 1.1Wan 2.7
Native audio + lip-syncYes — joint audio-video in a single pass; lip-sync in 7 languagesAudio-to-video listed among its modes; not positioned as audio-native
Clip duration3–15 seconds2–15 seconds
Output resolutionUp to 1080p, five aspect ratios720p or 1080p
Architecture15B unified self-attention Transformer (community-documented; no official report)27B total / 14B active MoE (published)
Open weightsNoYes for base models — Apache 2.0
Generation modes4 — text-, image-, subject-to-video + video editing (new in 1.1)7 — incl. start–end animation, video continue, audio-to-video, multi-reference
fal.ai price (per second)$0.14 (720p) · $0.28 (1080p)$0.10 flat — 2.8× cheaper at 1080p
Hedra-listed price (1080p, Jul 17)$9.90/min ≈ $0.165/sec$9.00/min ≈ $0.15/sec
Best fitDialogue-led hero spots where sync quality carries the creativeVolume output, edit-heavy revision cycles, budget-sensitive batches

HappyHorse resolution, clip-length, aspect-ratio, lip-sync and architecture figures are the published 1.0 spec, which Alibaba carries into 1.1 except where its 1.1 notes state otherwise; no standalone 1.1 spec sheet has been published. Wan 2.7 figures are from Alibaba's own published materials.

Hero creative
Dialogue-led spots & talking-head ads

HappyHorse's single-pass joint audio-video with seven-language lip-sync is the differentiator no rival on this stack matches. Pay the premium for final renders where sync quality is the creative.

Pick HappyHorse 1.1
Volume production
Drafts, variants & revisions

At fal's rates, $0.10/sec flat plus natural-language editing of nearly every element makes Wan 2.7 the iteration engine — roughly 2.8 seconds of draft per hero-render second at 1080p.

Pick Wan 2.7
Self-host / fine-tune
Brand-tuned or sovereignty-bound pipelines

Only one of the two lines has downloadable weights. Wan's Apache 2.0 base models are the sole option for on-prem deployment or fine-tuning on proprietary brand assets.

Wan 2.7 — only option
Procurement review
Enterprise & compliance context

The Pentagon added Alibaba to its Chinese-military-company list on June 8, 2026 — no direct sanctions, but added procurement friction for Western and defense-adjacent buyers. Alibaba said it is 'not a Chinese military company nor part of any military-civil fusion strategy.'

Run compliance early

The practical pattern for a marketing team is a pipeline, not a pick: draft and iterate on Wan 2.7, promote the surviving concepts to HappyHorse 1.1 for the final audio-native render, and re-quote host pricing each cycle. That is the same tiered-model discipline we apply when building AI-assisted social video production for clients, and it is the kind of vendor-routing decision our AI transformation engagements formalize — cheap model for volume, premium model where the capability difference is visible in the output.

09ConclusionA stack decision, not a launch chase.

The shape of Alibaba video, July 2026

Two models, two postures, one usable production stack.

Alibaba's video stack rewards being read as a system. HappyHorse 1.1 is the audio-native flagship with the strongest single differentiator in the lineup — joint audio-video generation with multi-language lip-sync in one pass — and Wan 2.7 is the workhorse: cheaper on every host that lists both, more editable, and the only line with open weights. Used together, they form a two-tier pipeline that neither forms alone.

The honest caveats matter as much as the specs. Arena standing is a dated snapshot, not a property — 1,444 on June 22 and 1,149 on July 17 are different platforms' reads, not a decline curve. Pricing is per-host, not per-model. And the company's openness posture varies so much by product line that every release needs its own verification rather than an assumption inherited from the Qwen brand.

The forward question is whether the sealed HappyHorse posture is a competitive window or a permanent stance. Either way, the buying advice holds: run your own evals on your own briefs, date-stamp every leaderboard figure you rely on, and price the pipeline — draft tier plus hero tier — rather than either model in isolation.

Put AI video into production

Tiered model pipelines make AI video economically sane.

Our team helps brands build tiered AI video pipelines — model selection, host pricing, draft-to-hero workflows, and brand-safety review — delivered in days, not quarters.

Free consultationExpert guidanceTailored solutions
What we work on

AI video production engagements

  • Model evals on your own creative briefs
  • Draft-tier / hero-tier pipeline design
  • Host and per-second pricing audits
  • Lip-sync and multi-language ad localization
  • Brand-safety and compliance review for AI video
FAQ · Alibaba video stack

The questions we get every week.

HappyHorse 1.1 is the current version of Alibaba's flagship AI video generation model, released June 21–22, 2026 — Alibaba Cloud's announcement post landed Sunday June 21, with trade coverage following through June 23. It is billed as a major upgrade over HappyHorse 1.0, with vendor-stated improvements in motion dynamics, subject consistency, prompt adherence, visual quality, and audio generation, plus a newly added video-editing mode. Its signature capability is joint audio-video generation in a single pass, including native lip-sync across seven languages. It is available on the HappyHorse website, Alibaba Cloud's Bailian platform, and Qwen Cloud, with enterprise API access via Alibaba Cloud Model Studio.
Related dispatches

Continue exploring the video field.