Back to Skills

p-video-avatar

Generate talking head avatar videos with Pruna P-Video-Avatar via inference.sh CLI. Turn a portrait image into a realistic speaking video with built-in TTS. 18x faster and 6x cheaper than competitors. Models: P-Video-Avatar, P-Image (for portrait generation). Capabilities: text-to-avatar, audio-driven avatars, 30 voices, 10 languages, 720p/1080p, built-in TTS, dynamic backgrounds, full-body control. Use for: AI presenters, product demos, explainer videos, virtual influencers, marketing, educatio

566stars88forksUpdated 6/26/2026

Security Assessment

Safe(100/100)
Security Score100/100

About p-video-avatar

P-Video-Avatar generates talking-head avatar videos from a single portrait image using Pruna's P-Video-Avatar model via the inference.sh CLI (the `belt` command). It turns a portrait into a realistic speaking video with built-in text-to-speech, positioned as the fastest and most cost-effective avatar model available — quality on par with Veo 3.0, and per the doc roughly 18x faster and 6x cheaper than alternatives such as Fabric, OmniHuman, and HeyGen. It requires the belt CLI and starts with `belt login`. A full workflow is supported where Pruna P-Image first generates a portrait (suggested aspect ratio 9:16 for vertical video) and P-Video-Avatar then animates it.

Generation uses `belt app run pruna/p-video-avatar --input '{...}'`. The required `image` parameter is a portrait (jpg, jpeg, png, webp). For speech, either provide `voice_script` text with a `voice` selection and `voice_language`, or supply your own `audio` file — and when both audio and voice_script are present, audio takes priority. Additional parameters include `resolution` (720p or 1080p), `video_prompt` to control avatar behavior and background, `voice_prompt` to control tone, pacing, and emotion, `seed` for reproducibility, `disable_safety_filter`, and `disable_prompt_upsampling`. The output video's aspect ratio matches the input image.

It ships 30 voices split across female and male sets and supports 10 languages: English (US), English (UK), Spanish, French, German, Italian, Portuguese (Brazil), Japanese, Korean, and Hindi. Pricing is $0.025 per second of output at 720p and $0.045 at 1080p (e.g. a 30-second 720p video is $0.75), with a free launch weekend noted from May 1 to May 4, 2026 (CET). Use cases span marketing demos and UGC ads, education and explainers, multilingual localization, virtual influencers, corporate training, gaming avatars, and personalized support videos. Tips recommend high-quality front-facing portraits, keeping videos under three minutes for visual consistency, and using video_prompt and voice_prompt for finer control. The skill is scoped to the `Bash(belt *)` tool.

FAQ

What input does it need to create an avatar video?

A portrait image (jpg, jpeg, png, or webp) is required. You then provide either a voice_script with a chosen voice and language, or your own audio file.

What happens if I provide both a voice_script and an audio file?

Audio takes priority. When both audio and voice_script are supplied, the avatar speaks from the audio file.

How many voices and languages are available?

There are 30 voices split between female and male sets, and 10 supported languages including English (US), English (UK), Spanish, French, German, Italian, Portuguese (Brazil), Japanese, Korean, and Hindi.

How much does it cost?

Output is billed at $0.025 per second at 720p and $0.045 per second at 1080p, so a 30-second 720p video costs $0.75. A free launch weekend ran from May 1 through May 4, 2026 (CET).

How do I control the avatar's background, tone, and pacing?

Use video_prompt to control avatar behavior and background, and voice_prompt to control tone, pacing, and emotion. The output aspect ratio matches the input portrait image.

Install p-video-avatar

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill