Skip to content
Clipia.
Home
What are we creating?

Pick a mode — Studio opens with the right prompt ready

All templates
Text
AI AssistantChat, ideas and scriptsPrompt EngineerCraft the perfect promptConcept DesignerVibe, palette, moodboard and refsTutorStep-by-step, explains the why
Images
Text → imageGenerate an image from a promptEditModify an existing photoImage templatesReady-made product shots and artAI Image GeneratorImages from text or photos
Video
Text → videoDescribe a scene and get a clipImage → videoBring a frame to life with motionVideo templatesReady-made scenes and stylesAI Video GeneratorText and image to finished videoPhoto → Video AIAnimate a frame with AIStudioVideo and images in one flow
Pricing
AboutTeam, mission and contactsWhat's NewLatest features and updatesBlogGuides, cases and newsModelsAI model catalogSupportTech help and contacts
For BusinessAPI, MCP and transparent billingSolutions by industryAI for interiors, ads, restaurants and more nichesFor PartnersReferral program
MCP & CLINewGenerate in Claude, Cursor & CLIAI GatewayOpenAI-compatible LLM APIDocumentationAPI reference and guidesDeveloper APIREST API and SDK
Sign In
  • Home

  • Assistant

    NEW
  • Create Video

  • Create Image

  • Studio

  • Turnkey video

  • My Works

  • Models

  • Support

  1. Home/
  2. Video Models/
  3. HappyHorse 1.0

Ready to create the impossible?

Welcome credits are already on your account — start creating.

Clipia.

Think differently — create the impossible.

All systems operational

Creation

  • Create Image
  • Create Video
  • AI Video Generator
  • AI Image Generator
  • Photo to Video AI
  • Templates

Solutions

  • For interiors
  • For advertising
  • For ad videos
  • For restaurants
  • For online schools
  • For game dev
  • For events
  • For presentations
  • All solutions

Models & resources

  • AI Models
  • Claude Opus 5
  • Gemini 3.6 Flash
  • Video Models
  • Image Models
  • Model Rankings
  • What's newNEW

Help

  • Guides
  • Documentation
  • About
  • Contact Us
  • Telegram Support

Legal

  • Terms of Service
  • Privacy Policy
  • Payment & Refund
  • Refund Policy
  • Content License
  • Cross-Border Transfers
  • Partner Agreement
© 2026 Clipia.ai. All rights reserved.
CookiesAcceptable Use

Clipia is an independent platform. Third-party AI model names are used solely to identify integrations available through our interface and do not imply affiliation with or endorsement by their respective owners.

HappyHorse 1.0 — Alibaba ATH

HappyHorse 1.0 video #1 on Video Arena

Alibaba's video model — a unified 15-billion-parameter Transformer. Native audio in a single pass, lip-sync in 7 languages, up to 15 seconds at 1080p. Four modes: T2V, I2V, Reference-to-Video and editing of an existing clip.

What it does
up to 15 secduration
1080pquality
from 30credits
Result

What HappyHorse 1.0 can do

Four modes inside one model — no tool switching

Video from text — with sound

Describe the scene in words — the model returns a clip up to 15 seconds long with synchronised audio and realistic motion physics. No separate TTS, no editing.

Bring a photo to life

Upload an image — the model adds motion, smooth camera, wind, lighting shifts and synced audio. The original style is preserved.

Hold subject identity

Pass 1–9 reference images of the hero — the model keeps their face, outfit and proportions consistent across every frame of the full 15 seconds. A mode no other model offers.

Rewrite an existing clip with text

Upload an MP4 (up to 100 MB, 3–60 sec) and describe the change: swap the background, change the time of day, add objects, restyle the scene. Edit cadence is preserved.

Audio and video — in a single pass

A unified Transformer without cross-attention emits the video stream and audio track synchronously. That removes the drift typical of two-stage TTS pipelines.

Lip-sync in 7 languages

Full mouth-shape sync for English, Mandarin, Cantonese, Japanese, Korean, German and French. Prompts can be written in any of them — up to 2,500 characters.

Get started in 60 seconds

No API setup — Clipia has HappyHorse 1.0 wired in

1

Open /create-video and pick HappyHorse 1.0

Four modes in one model: T2V for text, I2V for animating a photo, R2V for keeping a character, Edit for restyling an existing clip.

2

Describe the scene or upload an asset

Write a prompt up to 2,500 characters in any of the 7 supported languages. The soundscape is described directly in the prompt.

3

Choose settings

Pick a duration from 3 to 15 seconds, 720p or 1080p, and the aspect ratio you need (16:9, 9:16, 1:1, 4:3, 3:4).

4

Get a ready-to-use MP4

~38 seconds for a 5-second 1080p clip. Audio is baked into the file — no additional editing.

Use cases

Where HappyHorse 1.0 delivers the most value

Social content

Reels, Shorts and TikTok with built-in music and effects. 9:16, 15 seconds, no separate audio edit.

Ads and performance

Performance-grade video clips. Reference-to-Video keeps a brand character consistent across an entire creative series.

Character animation

Stills, portraits and illustrations come to life — the model adds motion while preserving the look and style of the source.

Product motion

Catalog shots turn into cinematic clips with smooth camera, lighting and material texture rendered in motion.

Storytelling and shorts

Short scenes with dialogue, music and ambient sound — the foundation for indie web series and podcast-style clips.

Global content

One script — multiple language tracks. Lip-sync in 7 languages without a separate voice-over or reshoot.

Frequently asked questions

Everything you need to know about HappyHorse 1.0

HappyHorse 1.0 is a video model from Alibaba ATH led by Zhang Di — the former technical architect of Kling AI. It is a unified 40-layer Transformer with 15 billion parameters that generates video and audio in one pass. The model first appeared anonymously on Artificial Analysis on April 7, 2026 and took #1 within 24 hours. On April 10 Bloomberg and CNBC revealed Alibaba as the author.

On the AA Arena blind test HappyHorse leads Seedance 2.0 in both no-audio categories: T2V 1366 vs 1273, I2V 1391 vs 1356. With audio it flips — Seedance 2.0 holds #1 (T2V 1219, I2V 1162) and HappyHorse sits at #2 (1204 / 1159). If you care about motion physics, optics and subject stability without a soundtrack, HappyHorse wins. If a native audio track is the priority, Seedance 2.0 is ahead by a small margin.

Three reasons. Architecturally, it is a unified Transformer with no cross-attention that emits video and audio in a single pass — something only closed flagship models used to do. On benchmarks, it took #1 on AA Arena anonymously within 24 hours of release, beating Seedance, Kling, PixVerse and SkyReels. On scope, four modes live inside one model (T2V/I2V/R2V/Edit) instead of four separate products.

On Artificial Analysis Video Arena, an independent blind-test leaderboard, HappyHorse holds #1 in both no-audio categories: Text-to-Video (ELO 1366) and Image-to-Video (ELO 1391). On motion physics, optics and subject stability the model is currently the strongest publicly available option. With audio it sits at #2.

Four modes inside one model: Text-to-Video (T2V) — from a prompt, Image-to-Video (I2V) — bring a still to life, Reference-to-Video (R2V) — hold subject identity across the entire clip, and Video Editing — restyle an existing clip with a text command. Every mode supports audio and lip-sync.

HappyHorse uses a unified Transformer with no cross-attention: the video stream and the audio track are generated jointly in a single forward pass. That removes the drift you tend to see with two-stage TTS pipelines. You can describe the soundscape directly in the prompt and the model will use it when scoring the clip.

Seven languages with full mouth-shape sync: English, Mandarin, Cantonese, Japanese, Korean, German and French. Prompts can be written in any of them — up to 2,500 characters.

720p and 1080p. Five aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4. I2V accepts JPEG, PNG and WEBP up to 10 MB; Video Editing accepts MP4 or MOV up to 100 MB and 3–60 seconds long.

From 3 to 15 seconds, 1-second steps (any integer). Average generation time is roughly 38 seconds for a 5-second 1080p clip. For longer narratives you can chain several clips with the same character through R2V.

Pricing starts from 30 credits and depends on duration and resolution. The exact amount is shown before generation starts. Audio and lip-sync are included — there is no separate fee for the soundtrack.

Generate the world's #1 video on Clipia

HappyHorse 1.0 — one model for T2V, I2V, Reference-to-Video and editing. With native audio and lip-sync in 7 languages.