digital-health-clinical-asr-eval
Stage 3 of Clinical ASR Flywheel. Score a NeMo manifest, produce the five-section KER leaderboard (by-ipa_source diagnostic). Not for ASR auth (/riva-asr).
Security Assessment
About digital-health-clinical-asr-eval
This skill is Stage 3 of NVIDIA's Clinical ASR Flywheel: a score-and-route stage that takes a NeMo-format manifest, transcribes it through a hosted NVIDIA ASR NIM, scores four metrics (WER/CER/KER/SER), and produces a five-section KER leaderboard with a by-ipa_source diagnostic. It solves the problem of objectively measuring clinical automatic-speech-recognition quality on domain-specific terminology, then deciding whether to advance to fine-tuning, loop back to data generation, or stop and harden the evaluation.
The workflow surfaces a clear data-transmission disclosure before any audio is sent: each manifest row's WAV plus its reference text goes to an external NVIDIA NVCF ASR service, while scoring runs locally in pure Python (or jiwer if installed). It strongly instructs users to pass only synthetic audio generated by Stage 2 and never real patient encounters or PHI. It defaults to the Parakeet TDT 0.6B v2 offline NIM with documented env-var overrides for model name, function ID, and self-hosted endpoints, and echoes the resolved function ID before spending API credits. A companion reference carries the full gRPC recipe and an alternate-NIM catalogue, and the skill routes off-topic requests (auth, model selection, deployment) to sibling skills. It requires an NVIDIA_API_KEY and treats its own SKILL.md as self-contained for methodology questions.
It targets healthcare ML engineers and TME teams building and evaluating clinical ASR models. Typical use cases include benchmarking a candidate ASR NIM against a curated clinical term list, generating a leaderboard to compare Parakeet versus Whisper, and gating progression through the flywheel.
FAQ
What does this skill produce?
A five-section KER leaderboard with a by-ipa_source diagnostic, computed from four scored metrics (WER, CER, KER, SER) over a NeMo-format manifest, plus a decision-tree recommendation to advance, loop back, or harden the eval.
What data leaves my environment?
During the ASR step, each manifest row's WAV bytes plus the reference transcript and clinical-extension metadata are sent to NVIDIA's hosted NVCF ASR service. Scoring itself runs locally and transmits nothing. The skill discloses this before the first call.
Can I use real patient audio?
No. The skill explicitly instructs users to pass only synthetic audio generated by Stage 2 and never to send real ASR recordings, real patient encounters, or any PHI through it.
What are the prerequisites?
An NVIDIA_API_KEY for hosted ASR NIMs via NVCF and a NeMo-format manifest (from the Stage 2 build skill or an externally provided one carrying the clinical-extension fields).
Which ASR model does it use and can I change it?
It defaults to nvidia/parakeet-tdt-0.6b-v2 (offline gRPC) but supports env-var overrides: ASR_MODEL_NAME for the display name, ASR_NVCF_FUNCTION_ID to swap to another hosted NIM like Whisper Large v3, and ASR_ENDPOINT for a self-hosted gRPC server.
Install digital-health-clinical-asr-eval
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
nvidia/skills