Alibaba's AI video stack is easy to misread as one product, but it is two distinct model lines run in parallel: HappyHorse, the audio-native flagship that appeared anonymously at the top of the Artificial Analysis Video Arena in April 2026, and Wan 2.7, the cheaper, more editable suite from Tongyi Lab with Apache 2.0 open weights for its base models. Neither is new — HappyHorse 1.1 shipped June 21–22 and the Wan 2.7 line launched April 6 — which is exactly why a stack-level read is now more useful than another launch recap.
The stakes for creative-production teams are practical, not academic. The two lines price differently on the same host, expose different generation modes, and sit at opposite ends of Alibaba's openness spectrum. Meanwhile the competitive field around them has contracted — VentureBeat's June analysis framed OpenAI's discontinued standalone Sora product and ByteDance's shelved international Seedance rollout as the backdrop against which Alibaba's stack now competes.
This guide covers how HappyHorse went from anonymous arena entry to claimed flagship, what 1.1 actually ships, how Wan 2.7 differs and why it costs less, why the two published arena Elo snapshots do not reconcile, how pricing varies by host, and — the part nobody else is writing — why the same company ships three different openness postures across three modalities.
- 01Two parallel video lines, neither of them new.HappyHorse 1.1 shipped June 21–22, 2026; the Wan 2.7 suite launched April 6 with incremental snapshots through June 12. Neither is launch news as of late July — treat this as a stack decision, not a release chase.
- 02HappyHorse debuted anonymously — and won.An unbranded model took the No. 1 spot on both text-to-video and image-to-video blind-test leaderboards around April 7, 2026. Alibaba confirmed authorship three days later, on April 10, per CNBC.
- 03Arena standing has drifted, and the snapshots don't reconcile.VentureBeat, citing Arena.ai on June 22, reported HappyHorse at Elo 1,444 and No. 2 across leaderboards. A July 17 Hedra read lists 1,149 T2V / 1,111 I2V. Different platforms, dates, and vote pools — never merge them into one figure.
- 04Pricing varies by host — roughly $0.10 to $0.28 per second.On fal.ai, Wan 2.7 runs $0.10/sec flat while HappyHorse runs $0.14/sec at 720p and $0.28/sec at 1080p. Hedra's July comparison lists different per-minute rates. There is no single canonical price.
- 05Alibaba's openness posture is inconsistent by design or by accident.Wan ships Apache 2.0 weights and published architecture. HappyHorse ships neither weights nor a tech report. Qwen-Image-3.0 went invite-only with no benchmarks. Three postures, one company, one April-to-July window.
01 — The StackOne company, two video lines.
Early speculation on X held that HappyHorse might secretly be Wan 2.7 under a codename. It is not — they are separate models from separate teams, and comparison leaderboards list them as separate entries. HappyHorse is the closed flagship; Wan 2.7 is Tongyi Lab's suite, with open weights for its base models and a published architecture. Understanding them as a two-tier stack — audio-native hero model on top, cheaper editable workhorse below — is the single most useful framing for a production team.
The field they compete in has thinned. For the wider post-Sora landscape and how Google's and xAI's entries stack up, see our AI video field comparison; for the July snapshot of FLUX 3, Seedance 2.5, and Gemini's video line, the FLUX 3-led video field comparison published alongside this post covers the non-Alibaba half of the market.
HappyHorse 1.1
Joint audio-video generation in a single pass, lip-sync across seven languages, four generation modes. No open weights, no official tech report — the 15B unified-Transformer architecture is community-documented.
Wan 2.7
27B total / 14B active MoE, Apache 2.0 license for base models, four sub-models and seven generation modes including natural-language video editing. Cheapest per second of the two on the same host.
02 — Origin StoryThe anonymous arena debut.
HappyHorse's origin is the most unusual go-to-market in the 2026 video-model field. Around April 7, 2026, an unbranded model calling itself "HappyHorse" appeared on the Artificial Analysis Video Arena — the blind-test platform where voters compare clips without knowing which model produced them — and immediately took the No. 1 spot on both the text-to-video and image-to-video leaderboards, ahead of rivals including ByteDance's Seedance line. Three days later, on April 10, CNBC reported Alibaba's confirmation that the model was its own.
Attribution below the company level is thinner. Trade press reports place the model with an Alibaba Token Hub (ATH) team with roots in the Taobao/Tmall Future Life Lab, restructured into a standalone unit — though outlets word the org chart differently, and Alibaba has not published its own account. The same trade coverage names Zhang Di, reported as a roughly 15-year AI veteran, former Kuaishou VP, and technical architect of Kling AI who rejoined Alibaba in late 2025, as the team's lead. Neither the exact lab name nor the team-lead attribution appears in CNBC's or VentureBeat's primary reporting, so treat both as trade-press-reported rather than Alibaba-confirmed.
The anonymous-debut strategy matters beyond trivia. Winning a blind test before revealing the brand neutralized the discount that Western evaluators sometimes apply to Chinese-lab releases — voters ranked the output first and learned the vendor afterward. It is also, notably, the inverse of how Alibaba launched Qwen-Image-3.0 three months later: maximum benchmark transparency first versus no benchmarks at all. That inconsistency is Section 07's subject.
Anonymous No. 1
An unbranded 'HappyHorse' entry took the top spot on both text-to-video and image-to-video blind-test leaderboards on the Artificial Analysis Video Arena, per CNBC's reporting.
Alibaba confirms
On April 10, 2026, Alibaba confirmed it built HappyHorse — three days after the anonymous debut. The blind-test-first strategy meant voters ranked output before knowing the vendor.
fal.ai goes live
fal launched developer and enterprise access to HappyHorse-1.0 as the first official third-party API partner, with four endpoints: text-to-video, image-to-video, reference-to-video, and video-edit.
03 — The FlagshipWhat HappyHorse 1.1 actually ships.
HappyHorse 1.1 arrived June 21–22, 2026 — Alibaba Cloud's own post landed Sunday June 21, with TechNode's write-up following on June 23. Alibaba billed it as a major upgrade over 1.0, with vendor-stated improvements in motion dynamics, subject consistency, prompt adherence, visual quality, and audio generation. No independent before/after benchmark has been published, so those deltas remain the vendor's own claims.
The mode lineup grew to four: text-to-video, image-to-video, subject-to-video, and — new in 1.1 — video editing, all with integrated audio at no additional cost through a single API endpoint. The carried-over 1.0 spec is the differentiator: 1080p output, 3–15 second clips, five aspect ratios, and native lip-sync across seven languages (Mandarin, Cantonese, English, Japanese, Korean, German, French). Architecturally, HappyHorse generates audio and video jointly in a single pass through a unified 40-layer self-attention Transformer with no cross-attention modules — per fal's product documentation — rather than stitching a video model to a separate audio model. Secondary trade coverage additionally describes synchronized dialogue, ambient sound, and effects in one inference step at a reported 24fps, though those specifics do not appear verbatim on Alibaba's or fal's own pages.
The parameter figure deserves its own hedge: the widely cited 15-billion-parameter unified self-attention Transformer — one token sequence carrying text, image, video, and audio tokens — comes from community-compiled technical documentation referenced by VentureBeat, because Alibaba has published no official HappyHorse paper. Availability is broad regardless: the HappyHorse website, Alibaba Cloud's Bailian platform, and Qwen Cloud for end users, with full API access via Alibaba Cloud Model Studio for enterprise customers. A 40% sitewide discount ran for the first two weeks after the 1.1 release — expired by this writing.
Video editing is new
Text-to-video, image-to-video, subject-to-video, and video editing — the last added in 1.1. All four carry integrated audio at no extra cost via one API endpoint.
Joint audio-video
Mandarin, Cantonese, English, Japanese, Korean, German, French — generated in the same single pass as the video rather than dubbed on afterward.
1080p, five ratios
Output up to 1080p in 16:9, 9:16, 1:1, 4:3, and 3:4 — the short-form envelope that covers most paid-social and organic placements without cropping.
04 — The WorkhorseWan 2.7: the open workhorse.
Wan 2.7 is the older, better-documented, and cheaper half of the stack. Tongyi Lab launched the full suite on April 6, 2026 — the image component surfaced around April 1 and text-to-video was live by April 3 — with incremental snapshot updates continuing through June 12. Where HappyHorse ships as a sealed product, Wan ships as an architecture: 27 billion total parameters with 14 billion active per step via mixture-of-experts, under an Apache 2.0 license, split across four sub-models — text-to-video, image-to-video, reference-to-video, and video edit.
The capability envelope trades audio-native polish for breadth. Clips run 2–15 seconds at 720p or 1080p across seven generation modes, including start–end frame animation, video continuation, audio-to-video, and multi-reference consistency. The most production-relevant feature is natural-language editing of nearly every video element — character actions, dialogue, appearance, scene, style, camera — which turns revision rounds from re-rolls into edits. For iteration-heavy marketing work, that editing surface plus the lower per-second price is often worth more than HappyHorse's superior sync.
05 — BenchmarksTwo Elo snapshots that don't reconcile.
Most coverage repeats one of two stories in isolation: the April "anonymous No. 1" story or the June "No. 2" story. Put the published numbers side by side, though, and an honest caveat emerges — the two most-cited Elo snapshots cannot be stitched into a single trend line, and anyone using "arena rank" in a buying decision should understand why.
No. 2 across leaderboards
VentureBeat, citing Arena.ai, reported HappyHorse at 1,444 Elo in both text-to-video and image-to-video — leading Google's Veo-3.1 (with audio) by 69 points in T2V and xAI's Grok-Imagine-Video by 23 points in I2V.
HappyHorse 1.1 · T2V
Hedra's model-comparison post lists HappyHorse 1.1 at 1,149 T2V / 1,111 I2V on 3,535 votes, calling it a serious #5 alternative at #4 overall and #4 on text-to-video. A different platform read, weeks later — not the June number minus decline.
Wan 2.7 · No. 3 T2V
The same Hedra read places Wan 2.7 at 1,159 Elo, No. 3 on the text-to-video leaderboard on 3,153 votes — the cheaper sibling out-ranking the flagship on that particular snapshot.
The 1,444 and 1,149 figures are not directly comparable. Arena Elo drifts as new models enter and vote pools accumulate, and the two sources may be reading different leaderboard categories at different moments. What the pair of dated snapshots supports is a qualitative claim: HappyHorse's standing became more contested between late June and mid-July as Veo 3.1 and other rivals closed the gap, and by Hedra's July read Wan 2.7 ranked above its own flagship sibling on text-to-video.
The interpretive lesson generalizes past Alibaba: arena rank is a moving, provider-dependent measurement, not a fixed property of a model. A leaderboard position quoted without a date and source is close to meaningless for procurement. If a video model's rank is load-bearing in your vendor decision, capture the snapshot — platform, date, category, vote count — the way you would any other perishable benchmark, and re-check it before renewal.
06 — EconomicsPricing varies by host, not just by model.
There is no single canonical price for either model — rates are set independently by each hosting provider. On fal.ai, HappyHorse's first official third-party API partner, HappyHorse runs $0.14 per second at 720p and $0.28 per second at 1080p, while Wan 2.7 runs a flat $0.10 per second at either resolution — a ten-second 1080p clip costs $1.00 on Wan versus $2.80 on HappyHorse, a 2.8× gap that narrows to 1.4× at 720p. Hedra's July 17 comparison lists different rates again: $9.00 per minute for Wan 2.7 at 1080p (≈$0.15/sec) and $9.90 per minute for HappyHorse 1.1 (≈$0.165/sec). All rates are pay-per-use with no minimums on fal.
Per-second cost · Alibaba video models across hosts
Sources: fal.ai product pages (retrieved Jul 24, 2026); Hedra model-comparison blog (Jul 17, 2026). Rates set independently per host.Two readings follow. First, the spread across hosts — HappyHorse 1080p at $0.28/sec on fal versus roughly $0.165/sec in Hedra's listing — means there is no single canonical rate for the same model, which argues for re-quoting rates at the start of every production cycle rather than baking one listing's number into a budget. (Hedra publishes a comparison; its role as a host is not confirmed, so treat its figures as a listing, not a competing provider's price.) Second, the model gap on a single host is large enough to shape workflow: at fal's rates, every hero-quality HappyHorse second at 1080p buys 2.8 seconds of Wan output. Teams that iterate — which is every team — can afford nearly three times the draft volume on the workhorse before committing the flagship to final renders.
07 — The PatternThree modalities, three openness postures.
Here is the angle the launch coverage misses: between April and July 2026, the same company shipped three flagship-tier generative lines with three incompatible openness postures. Wan 2.7 shipped with Apache 2.0 weights for its base models and a published architecture. HappyHorse shipped with no weights and no tech report — its architecture exists publicly only as community-compiled documentation. And Qwen-Image-3.0, launched July 21, went further still: API and invite-only, with no weights, no benchmarks, no license, and no tech report, a sharp break from Qwen-Image 1.0's Apache 2.0 release with a published report in August 2025.
Open posture
27B/14B MoE detailed publicly, base-model weights downloadable, four sub-models documented. The posture that built Alibaba's open-model credibility.
Sealed posture
Benchmark-transparent (it debuted on a public blind-test arena) but artifact-closed: the 15B unified-Transformer description exists only as community documentation.
Closed posture
Launched Jul 21, 2026 with striking capabilities — 4,500-token prompts, legible 10px text — and zero published evidence, reversing Qwen-Image 1.0's open release.
We covered the image-side retreat in detail in Alibaba's broader closed-flagship pivot and Qwen-Image-3.0's own closed pivot. Read against the video stack, the pattern looks less like a single strategy shift and more like per-team autonomy: the Wan and Qwen open-weight tradition on one side, and newer flagship-chasing units on the other, each choosing its own disclosure level. For buyers, the practical consequence is that "Alibaba is an open-model company" is no longer a safe planning assumption — openness now has to be verified per product line, per release.
Looking forward, two signals are worth watching. First, whether third-party hosts that carried HappyHorse 1.0 update to 1.1 — VentureBeat explicitly framed that as an open indicator of developer demand as of June 22. Second, whether HappyHorse ever gets a tech report or weights: if it does, the sealed posture was a competitive-window tactic; if it doesn't, Alibaba's video stack settles permanently into a closed-flagship-plus-open-workhorse shape — which, notably, is the same two-tier structure most Western labs have converged on.
08 — Decision MatrixWhich model for which job.
No swept source puts both hosts' pricing and both models' specs side by side — coverage treats HappyHorse origin stories and Wan spec sheets as separate beats. The table below is that missing head-to-head, compiled from fal.ai's product pages, Hedra's July comparison, TechNode, and VentureBeat.
| Spec | HappyHorse 1.1 | Wan 2.7 |
|---|---|---|
| Native audio + lip-sync | Yes — joint audio-video in a single pass; lip-sync in 7 languages | Audio-to-video listed among its modes; not positioned as audio-native |
| Clip duration | 3–15 seconds | 2–15 seconds |
| Output resolution | Up to 1080p, five aspect ratios | 720p or 1080p |
| Architecture | 15B unified self-attention Transformer (community-documented; no official report) | 27B total / 14B active MoE (published) |
| Open weights | No | Yes for base models — Apache 2.0 |
| Generation modes | 4 — text-, image-, subject-to-video + video editing (new in 1.1) | 7 — incl. start–end animation, video continue, audio-to-video, multi-reference |
| fal.ai price (per second) | $0.14 (720p) · $0.28 (1080p) | $0.10 flat — 2.8× cheaper at 1080p |
| Hedra-listed price (1080p, Jul 17) | $9.90/min ≈ $0.165/sec | $9.00/min ≈ $0.15/sec |
| Best fit | Dialogue-led hero spots where sync quality carries the creative | Volume output, edit-heavy revision cycles, budget-sensitive batches |
HappyHorse resolution, clip-length, aspect-ratio, lip-sync and architecture figures are the published 1.0 spec, which Alibaba carries into 1.1 except where its 1.1 notes state otherwise; no standalone 1.1 spec sheet has been published. Wan 2.7 figures are from Alibaba's own published materials.
Dialogue-led spots & talking-head ads
HappyHorse's single-pass joint audio-video with seven-language lip-sync is the differentiator no rival on this stack matches. Pay the premium for final renders where sync quality is the creative.
Drafts, variants & revisions
At fal's rates, $0.10/sec flat plus natural-language editing of nearly every element makes Wan 2.7 the iteration engine — roughly 2.8 seconds of draft per hero-render second at 1080p.
Brand-tuned or sovereignty-bound pipelines
Only one of the two lines has downloadable weights. Wan's Apache 2.0 base models are the sole option for on-prem deployment or fine-tuning on proprietary brand assets.
Enterprise & compliance context
The Pentagon added Alibaba to its Chinese-military-company list on June 8, 2026 — no direct sanctions, but added procurement friction for Western and defense-adjacent buyers. Alibaba said it is 'not a Chinese military company nor part of any military-civil fusion strategy.'
The practical pattern for a marketing team is a pipeline, not a pick: draft and iterate on Wan 2.7, promote the surviving concepts to HappyHorse 1.1 for the final audio-native render, and re-quote host pricing each cycle. That is the same tiered-model discipline we apply when building AI-assisted social video production for clients, and it is the kind of vendor-routing decision our AI transformation engagements formalize — cheap model for volume, premium model where the capability difference is visible in the output.
09 — ConclusionA stack decision, not a launch chase.
Two models, two postures, one usable production stack.
Alibaba's video stack rewards being read as a system. HappyHorse 1.1 is the audio-native flagship with the strongest single differentiator in the lineup — joint audio-video generation with multi-language lip-sync in one pass — and Wan 2.7 is the workhorse: cheaper on every host that lists both, more editable, and the only line with open weights. Used together, they form a two-tier pipeline that neither forms alone.
The honest caveats matter as much as the specs. Arena standing is a dated snapshot, not a property — 1,444 on June 22 and 1,149 on July 17 are different platforms' reads, not a decline curve. Pricing is per-host, not per-model. And the company's openness posture varies so much by product line that every release needs its own verification rather than an assumption inherited from the Qwen brand.
The forward question is whether the sealed HappyHorse posture is a competitive window or a permanent stance. Either way, the buying advice holds: run your own evals on your own briefs, date-stamp every leaderboard figure you rely on, and price the pipeline — draft tier plus hero tier — rather than either model in isolation.