Back to Skills

muapi-ugc-video-factory

Turn a person photo + a product photo + an optional script into a vertical 9:16 UGC-style video ad. Generates a lifestyle hero image (Nano-Banana Pro Edit), then animates it with native audio using Seedance 2.0 VIP image-to-video.

4,040stars458forksUpdated 8/13/2026

Security Assessment

Low Risk(88/100)

Detected risks:

Realistic Talking-Person Video Generation(The pipeline fuses a supplied person photo into a hero image and animates it so the depicted person appears to speak a scripted line with synced audio, which could be misused to make a real person appear to say things they did not without their consent., The concern is limited because inputs are user-supplied and the tool is framed for the user's own UGC product advertising, but it carries standard likeness and consent responsibilities.)
Security Score88/100

About muapi-ugc-video-factory

muapi-ugc-video-factory is a generative-media skill that turns a person photo plus a product photo (and an optional script and environment description) into a vertical 9:16 UGC-style video advertisement with native dialogue audio. It solves the practical marketing problem of producing an authentic, social-media-ready product ad quickly, without a shoot, by orchestrating a three-stage AI pipeline through the muapi CLI.

The workflow runs three sequential steps whose outputs chain together. First, a GPT chat model (temperature 0, ~200 tokens) writes a director-grade ultra-realistic lifestyle photography prompt from the person, product, and environment inputs. Second, the Nano-Banana Pro Edit model fuses the person and product reference images into a single 1K, 9:16 hero photo, which is shown to the user for approval. Third, Seedance 2.0 VIP image-to-video animates that hero photo into a 10-second vertical clip with synced spoken audio, using a defined CFG scale, negative prompt, and a UGC-style prompt built around the user's script. The skill documents input handling (uploading local files via muapi upload, generating placeholders if inputs are missing), tier notes (VIP supports 9:16 and 4-15s durations, tolerates realistic human faces, and offers a faster variant), and authentication via a MUAPI_API_KEY environment variable with a raw-API curl fallback to the muapi endpoint.

It targets marketers, e-commerce sellers, and content creators who want talking-head UGC ads from a single person and product image. A noteworthy consideration is that the pipeline generates a realistic person appearing to speak a scripted line, so it carries synthetic-media/likeness responsibilities: it should be used with imagery the user is authorized to use. Credential handling via an environment variable is standard, and there is no destructive or exfiltrating behavior.

FAQ

What inputs does it need?

A required person photo and product photo, plus an optional spoken script (defaults to a short sample line) and an optional environment/scene description. If the person or product image is missing, the skill asks the user to upload one or offers to generate placeholders.

What is the generation pipeline?

Three sequential stages: a GPT model writes a lifestyle-photography prompt, Nano-Banana Pro Edit fuses the person and product into a 1K 9:16 hero image (shown for approval), and Seedance 2.0 VIP image-to-video animates it into a 10-second vertical clip with native dialogue audio.

How is authentication handled?

Through the muapi CLI, which is configured with a MUAPI_API_KEY environment variable (run muapi auth configure first if it is unset). A raw-API fallback sends the key as an x-api-key header to the muapi endpoint. The key is read from the environment, not hardcoded.

What are the output constraints and tips?

VIP tier supports 9:16 and 4-15 second durations, with 10 seconds ideal for a 1-2 sentence script; longer scripts get compressed and words clipped. For multi-shot ads, generate several hero-image variations and animate each independently, since VIP does not do multi-image image-to-video at 9:16 with audio.

Are there responsible-use considerations?

Yes. The pipeline produces a realistic person appearing to speak a scripted line, so it should be used only with a person's likeness and product imagery the user is authorized to use; the tool itself is intended for UGC-style product advertising.

Install muapi-ugc-video-factory

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill