Video from text — with sound
Describe the scene in words — the model returns a clip up to 15 seconds long with synchronised audio and realistic motion physics. No separate TTS, no editing.
Alibaba's video model — a unified 15-billion-parameter Transformer. Native audio in a single pass, lip-sync in 7 languages, up to 15 seconds at 1080p. Four modes: T2V, I2V, Reference-to-Video and editing of an existing clip.
Four modes inside one model — no tool switching
Describe the scene in words — the model returns a clip up to 15 seconds long with synchronised audio and realistic motion physics. No separate TTS, no editing.
Upload an image — the model adds motion, smooth camera, wind, lighting shifts and synced audio. The original style is preserved.
Pass 1–9 reference images of the hero — the model keeps their face, outfit and proportions consistent across every frame of the full 15 seconds. A mode no other model offers.
Upload an MP4 (up to 100 MB, 3–60 sec) and describe the change: swap the background, change the time of day, add objects, restyle the scene. Edit cadence is preserved.
A unified Transformer without cross-attention emits the video stream and audio track synchronously. That removes the drift typical of two-stage TTS pipelines.
Full mouth-shape sync for English, Mandarin, Cantonese, Japanese, Korean, German and French. Prompts can be written in any of them — up to 2,500 characters.
No API setup — Clipia has HappyHorse 1.0 wired in
Four modes in one model: T2V for text, I2V for animating a photo, R2V for keeping a character, Edit for restyling an existing clip.
Write a prompt up to 2,500 characters in any of the 7 supported languages. The soundscape is described directly in the prompt.
Pick a duration from 3 to 15 seconds, 720p or 1080p, and the aspect ratio you need (16:9, 9:16, 1:1, 4:3, 3:4).
~38 seconds for a 5-second 1080p clip. Audio is baked into the file — no additional editing.
Where HappyHorse 1.0 delivers the most value
Reels, Shorts and TikTok with built-in music and effects. 9:16, 15 seconds, no separate audio edit.
Performance-grade video clips. Reference-to-Video keeps a brand character consistent across an entire creative series.
Stills, portraits and illustrations come to life — the model adds motion while preserving the look and style of the source.
Catalog shots turn into cinematic clips with smooth camera, lighting and material texture rendered in motion.
Short scenes with dialogue, music and ambient sound — the foundation for indie web series and podcast-style clips.
One script — multiple language tracks. Lip-sync in 7 languages without a separate voice-over or reshoot.
Everything you need to know about HappyHorse 1.0
HappyHorse 1.0 is a video model from Alibaba ATH led by Zhang Di — the former technical architect of Kling AI. It is a unified 40-layer Transformer with 15 billion parameters that generates video and audio in one pass. The model first appeared anonymously on Artificial Analysis on April 7, 2026 and took #1 within 24 hours. On April 10 Bloomberg and CNBC revealed Alibaba as the author.
On the AA Arena blind test HappyHorse leads Seedance 2.0 in both no-audio categories: T2V 1366 vs 1273, I2V 1391 vs 1356. With audio it flips — Seedance 2.0 holds #1 (T2V 1219, I2V 1162) and HappyHorse sits at #2 (1204 / 1159). If you care about motion physics, optics and subject stability without a soundtrack, HappyHorse wins. If a native audio track is the priority, Seedance 2.0 is ahead by a small margin.
Three reasons. Architecturally, it is a unified Transformer with no cross-attention that emits video and audio in a single pass — something only closed flagship models used to do. On benchmarks, it took #1 on AA Arena anonymously within 24 hours of release, beating Seedance, Kling, PixVerse and SkyReels. On scope, four modes live inside one model (T2V/I2V/R2V/Edit) instead of four separate products.
On Artificial Analysis Video Arena, an independent blind-test leaderboard, HappyHorse holds #1 in both no-audio categories: Text-to-Video (ELO 1366) and Image-to-Video (ELO 1391). On motion physics, optics and subject stability the model is currently the strongest publicly available option. With audio it sits at #2.
Four modes inside one model: Text-to-Video (T2V) — from a prompt, Image-to-Video (I2V) — bring a still to life, Reference-to-Video (R2V) — hold subject identity across the entire clip, and Video Editing — restyle an existing clip with a text command. Every mode supports audio and lip-sync.
HappyHorse uses a unified Transformer with no cross-attention: the video stream and the audio track are generated jointly in a single forward pass. That removes the drift you tend to see with two-stage TTS pipelines. You can describe the soundscape directly in the prompt and the model will use it when scoring the clip.
Seven languages with full mouth-shape sync: English, Mandarin, Cantonese, Japanese, Korean, German and French. Prompts can be written in any of them — up to 2,500 characters.
720p and 1080p. Five aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4. I2V accepts JPEG, PNG and WEBP up to 10 MB; Video Editing accepts MP4 or MOV up to 100 MB and 3–60 seconds long.
From 3 to 15 seconds, 1-second steps (any integer). Average generation time is roughly 38 seconds for a 5-second 1080p clip. For longer narratives you can chain several clips with the same character through R2V.
Pricing starts from 30 credits and depends on duration and resolution. The exact amount is shown before generation starts. Audio and lip-sync are included — there is no separate fee for the soundtrack.