happyhorse
Generate and edit videos with Alibaba HappyHorse 1.0 models via inference.sh CLI. Models: HappyHorse T2V, I2V, R2V, Video Edit. Capabilities: text-to-video, image-to-video, reference-to-video, video editing with natural language, character preservation, 720P/1080P, up to 15 seconds. Use for: physically realistic video, video editing, character-consistent content, product demos, social media. Triggers: happyhorse, happy horse, alibaba video, happyhorse 1.0, dashscope video, alibaba happyhorse, vi
Security Assessment
About happyhorse
happyhorse generates and edits physically realistic videos with Alibaba's HappyHorse 1.0 models through the inference.sh CLI (the `belt` command). It exposes four models: T2V (`alibaba/happyhorse-1-0-t2v`) for text-to-video with realistic motion, I2V (`alibaba/happyhorse-1-0-i2v`) to animate a single image, R2V (`alibaba/happyhorse-1-0-r2v`) to preserve characters from up to 9 reference images, and Video Edit (`alibaba/happyhorse-1-0-video-edit`) to edit existing videos with natural language. All models support 720P/1080P resolution and durations up to 15 seconds. It requires the belt CLI (installable via the belt CLI skill) and begins with `belt login`.
Generation uses `belt app run <app-id> --input '{...}'`. T2V takes a required `prompt` with optional `duration` (3-15, default 5), `resolution` (720P/1080P, default 720P), `ratio` (16:9, 9:16, 1:1, 4:3, 3:4, 21:9; default 16:9), `seed`, and `watermark`. I2V requires a `first_frame` image (JPEG, PNG, WebP) with an optional prompt. R2V requires a `prompt` and a `reference_images` array of up to 9 character images, supporting multi-character scenes. Video Edit requires a `video` (MP4/MOV, H.264) and an editing `prompt`, and optionally accepts up to 5 `reference_images` and an `audio_setting` of auto, generate, or keep_original — for example replacing a person with a referenced character or changing a scene to a rainy day.
Pricing is $0.14 per second at 720P and $0.24 per second at 1080P, and Video Edit is billed on input plus output duration. The documented examples cover slow-motion text-to-video, image animation, single- and multi-character reference-to-video, background changes, character replacement, and audio-controlled editing. Users can find HappyHorse apps with `belt app store search "happyhorse"` and browse all video apps with `belt app store --category video`. The skill links to documentation for running apps, streaming results, and a content-pipeline example, and references related skills for the full inference.sh platform, broader video generation, Seedance 2.0, Google Veo, and image generation. It is scoped to the `Bash(belt *)` tool.
FAQ
What are the four HappyHorse models for?
T2V for text-to-video with realistic motion, I2V to animate a single image, R2V to preserve characters from up to 9 reference images, and Video Edit to edit existing videos with natural language instructions.
What resolution and duration are supported?
All models support 720P and 1080P resolution and durations up to 15 seconds (the duration parameter accepts 3 to 15, defaulting to 5).
How much does generation cost?
$0.14 per second at 720P and $0.24 per second at 1080P. Video Edit is billed on combined input plus output duration.
How many reference images can I use to keep a character consistent?
R2V accepts up to 9 character reference images and supports multi-character scenes. The Video Edit model accepts up to 5 reference images.
Can I control audio when editing a video?
Yes. The Video Edit model has an audio_setting parameter that can be auto, generate, or keep_original.
Install happyhorse
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
halt-catch-fire/skills