vss-summarize-video
Use to summarize a recorded video via the LVS summarization microservice (HITL-gated) with a VLM fallback. Not for report generation or live RTSP captioning.
Security Assessment
About vss-summarize-video
VSS Summarize Video produces a single polished narrative summary of one recorded video clip, with timestamped events when NVIDIA's LVS summarization microservice is reachable, and a direct VLM NIM fallback when it is not. It solves the focused problem of turning a recorded clip into a scenario-aware summary, and it deliberately scopes itself narrowly: it is not for live RTSP captioning, incident or alert-window report generation, or semantic search across an archive, each of which is delegated to a sibling skill.
The skill drives the workflow by calling the summarization microservice or VLM endpoint directly with `curl` and `jq`, following reference-routed sections for the summarization API (endpoints such as `/v1/summarize`, health probes, request fields, and response shapes), deployment and service ops (the `lvs` profile, ports, env vars, backend and datastore selection across Elasticsearch/Neo4j/ArangoDB), environment variables, and debugging. A human-in-the-loop gate is central: before any summarize call the skill must collect `scenario`, `events`, and optional `objects_of_interest` from the user and wait for a response, applying generic defaults only on explicit opt-in. It also documents troubleshooting for warm-up 503s, empty results, Cosmos reasoning `<think>` blocks, and checking HTTP status rather than response bodies.
Target users are operators and developers running the NVIDIA Video Search and Summarization blueprint who need to summarize recorded clips reliably. It assumes network reachability to the endpoints, fetchable clip URLs, and `jq`/`curl` on the agent host, and it is explicit about limitations: the VLM fallback uses one fixed prompt with lower quality, remote endpoints often cannot reach localhost/private clip URLs, and each request makes a single backend call with no parallel hedging.
FAQ
What does this skill do and what should I not use it for?
It summarizes a single recorded video clip via the LVS microservice with a VLM fallback. Do not use it for live RTSP captioning, incident or alert-window report generation, or semantic archive search; those are handled by vss-deploy-dense-captioning, vss-generate-video-report Mode B, and vss-search-archive respectively.
What are the prerequisites?
A VSS lvs profile running on the host (port 38111) or a reachable VLM/RT-VLM fallback endpoint, network reachability from the agent host to both endpoints, fetchable clip URLs, and `jq` plus `curl` available on the agent host.
Why does it ask me for scenario and events before summarizing?
It is human-in-the-loop gated: before any summarize call it must collect scenario, events, and optional objects_of_interest and wait for your reply. It only applies generic defaults (activity monitoring / notable activity) when you explicitly opt in, so summaries target what you actually care about.
What are the known limitations?
The direct VLM fallback uses a single fixed prompt and cannot target scenario or events, so its output quality is lower than the LVS path. Remote VLM endpoints generally cannot reach localhost or private clip URLs, and each request makes one backend call with no parallel hedging or multi-pass summaries.
How do I debug a service that never becomes ready?
If `/v1/ready` returns 503 repeatedly the LVS service is likely still warming up; retry up to about 30 seconds. Always check the HTTP status code with curl rather than inspecting the body, since the service can legitimately return 200 with an empty body. Deeper diagnostics are in the debugging reference.
All Files
14 filesInstall vss-summarize-video
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
nvidia/skills