muapi-ugc-video-factory
Turn a person photo + a product photo + an optional script into a vertical 9:16 UGC-style video ad. Generates a lifestyle hero image (Nano-Banana Pro Edit), then animates it with native audio using Seedance 2.0 VIP image-to-video.
Security Assessment
Detected risks:
About muapi-ugc-video-factory
muapi-ugc-video-factory is a generative-media skill that turns a person photo plus a product photo (and an optional script and environment description) into a vertical 9:16 UGC-style video advertisement with native dialogue audio. It solves the practical marketing problem of producing an authentic, social-media-ready product ad quickly, without a shoot, by orchestrating a three-stage AI pipeline through the muapi CLI.
The workflow runs three sequential steps whose outputs chain together. First, a GPT chat model (temperature 0, ~200 tokens) writes a director-grade ultra-realistic lifestyle photography prompt from the person, product, and environment inputs. Second, the Nano-Banana Pro Edit model fuses the person and product reference images into a single 1K, 9:16 hero photo, which is shown to the user for approval. Third, Seedance 2.0 VIP image-to-video animates that hero photo into a 10-second vertical clip with synced spoken audio, using a defined CFG scale, negative prompt, and a UGC-style prompt built around the user's script. The skill documents input handling (uploading local files via muapi upload, generating placeholders if inputs are missing), tier notes (VIP supports 9:16 and 4-15s durations, tolerates realistic human faces, and offers a faster variant), and authentication via a MUAPI_API_KEY environment variable with a raw-API curl fallback to the muapi endpoint.
It targets marketers, e-commerce sellers, and content creators who want talking-head UGC ads from a single person and product image. A noteworthy consideration is that the pipeline generates a realistic person appearing to speak a scripted line, so it carries synthetic-media/likeness responsibilities: it should be used with imagery the user is authorized to use. Credential handling via an environment variable is standard, and there is no destructive or exfiltrating behavior.
FAQ
What inputs does it need?
A required person photo and product photo, plus an optional spoken script (defaults to a short sample line) and an optional environment/scene description. If the person or product image is missing, the skill asks the user to upload one or offers to generate placeholders.
What is the generation pipeline?
Three sequential stages: a GPT model writes a lifestyle-photography prompt, Nano-Banana Pro Edit fuses the person and product into a 1K 9:16 hero image (shown for approval), and Seedance 2.0 VIP image-to-video animates it into a 10-second vertical clip with native dialogue audio.
How is authentication handled?
Through the muapi CLI, which is configured with a MUAPI_API_KEY environment variable (run muapi auth configure first if it is unset). A raw-API fallback sends the key as an x-api-key header to the muapi endpoint. The key is read from the environment, not hardcoded.
What are the output constraints and tips?
VIP tier supports 9:16 and 4-15 second durations, with 10 seconds ideal for a 1-2 sentence script; longer scripts get compressed and words clipped. For multi-shot ads, generate several hero-image variations and animate each independently, since VIP does not do multi-image image-to-video at 9:16 with audio.
Are there responsible-use considerations?
Yes. The pipeline produces a realistic person appearing to speak a scripted line, so it should be used only with a person's likeness and product imagery the user is authorized to use; the tool itself is intended for UGC-style product advertising.
Install muapi-ugc-video-factory
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
samuraigpt/generative-media-skills