Stage 2 of the Clinical ASR Flywheel. Use when curating clinical terms, tagging IPA, and synthesizing a NeMo manifest. NOT for scoring (use /digital-health-clinical-asr-eval).
This is Stage 2 of NVIDIA's four-stage Clinical ASR Flywheel, the "curate-and-synthesize" step. It addresses the problem that generic word-error-rate metrics hide the clinically important failures of automatic speech recognition on specialized vocabulary such as drug names, procedures, anatomy, and conditions. The skill helps a user curate a clinical-specialty term list, tag pronunciations, synthesize evaluation audio through NVIDIA Magpie TTS, and write a NeMo-format manifest.jsonl (with clinical extension fields like term, entity_category, ipa_source, voice_id, noise_level, and context_type) that becomes the input to Stage 3 scoring against the headline KER (keyword error rate) metric.
The workflow is deliberately conversational and gated: it asks specialty-aware clarifying questions before proposing terms, walks the user through a two-tier IPA pipeline (manual override, then Merriam-Webster lookup, then Magpie G2P fallback), enforces a QA-mode audition gate before full Cartesian synthesis, and names KER as the metric users will see downstream. A pronunciation-pipeline reference provides a full Merriam-Webster respelling-to-IPA mapping table and two MW implementation paths: a stable JSON API path using DICTIONARY_API_KEY and a brittle HTML-scrape fallback. The documentation is notably security-conscious about API-key handling: never hard-code keys, prefer a secrets manager, keep HTTPS with SSL verification on, rotate on suspected exposure, and log the act of a call but never the key.
It is aimed at clinical AI/ML engineers, healthcare TME teams, and clinicians (or those working with them) building realistic ASR benchmarks. The skill repeatedly warns that curated content leaves the environment to Merriam-Webster and NVIDIA NVCF, and that no real patient records, transcripts, or PHI should be passed through any stage.
A NeMo-format manifest.jsonl tagged with clinical extension fields, plus the synthesized audio it references and a curated term seed CSV, all ready for scoring at the Stage 3 eval skill.
Stage 1 (the setup skill) must be completed first. NVIDIA_API_KEY is required for hosted Magpie TTS via NVCF; DICTIONARY_API_KEY is optional for Merriam-Webster Medical Dictionary lookup.
Curated terms are sent one HTTP request per term to Merriam-Webster, and generated clinical sentences (plus any SSML IPA wrappers) are sent to NVIDIA NVCF Magpie TTS on each synthesis call. The skill requires you to disclose this before invoking either service.
Yes. Leaving DICTIONARY_API_KEY unset and not running a scraper takes a path that skips MW and falls through to Magpie G2P, at the cost of weaker coverage on long-tail clinical terms.
No. The endpoints expect non-PHI synthetic content only. The skill explicitly instructs users not to pass real patient records, real ASR transcripts, or any PHI through it, and to confirm organizational data-governance policy if the term list itself is sensitive.
Quick Setup:
.claude/skills/Repository
nvidia/skills