Back to Skills

digital-health-clinical-asr-finetune

Stage 4 of the Clinical ASR Flywheel. Use when priority KER is above 0.3 to run stock NeMo SFT on Parakeet TDT v2 and offline cycle N+1 re-eval. NOT for generic word boosting (use /finetune-asr).

2,728stars318forksUpdated 7/30/2026

Security Assessment

Safe(92/100)
Security Score92/100

About digital-health-clinical-asr-finetune

This is Stage 4 of NVIDIA's Clinical ASR Flywheel, the "adapt-and-measure" step. It solves the problem of actually closing the improvement loop: once Stage 3 eval shows priority-category KER above 0.3 on a sufficiently large manifest, this skill runs stock NeMo supervised fine-tuning (SFT) on a Parakeet TDT v2 base model and re-evaluates offline as cycle N+1 to measure that the model quality genuinely improved. It reports an empirically verified result on a reference manifest of baseline KER 0.513 dropping to 0.128 after three epochs, a 75 percent relative reduction, with the largest gains on drug names.

The skill is self-contained and prescriptive: it supplies gate criteria (fire only when priority KER exceeds 0.3 and the manifest has at least 100 rows), a base-model selection table, a stock-NeMo-SFT recipe using /opt/NeMo/examples/asr/speech_to_text_finetune.py inside the nvcr.io/nvidia/nemo:25.11.01 container, a hyperparameter table with rationale, and a cycle-N+1 decision table. It carries explicit anti-footguns: do not use the broken adapter-mixin path on TDT/RNNT decoders, and do not fine-tune the streaming Nemotron base whose SFT path collapses. A deep-dive reference covers hyperparameter reasoning, when to stop tuning, and full Brev GPU provisioning. Notably, where a cloud install uses a curl-piped install script, the reference calls this an antipattern and prescribes mitigations: download first and verify with shasum before running, pin to a release tag, or use Homebrew on macOS.

Target users are clinical AI/ML and healthcare TME engineers with access to a CUDA host (24 GB VRAM comfortable, 16 GB workable) or a Brev cloud GPU. NVIDIA_API_KEY is required for the offline re-eval round-trip and any NIM deploy, and the finetune-asr and riva-asr-custom skills are expected alongside it.

FAQ

When should Stage 4 fine-tuning run?

Only when priority-category KER is above 0.3 and the manifest has at least 100 rows (at least 5 per priority category). Below those thresholds the skill routes back to the build stage to grow the manifest first.

What hardware and container are needed?

A CUDA host (24 GB VRAM comfortable, 16 GB workable with batch_size=4) and the NeMo container nvcr.io/nvidia/nemo:25.11.01. No local GPU users are pointed to Brev cloud instances. NVIDIA_API_KEY is required for the offline eval round-trip and any NIM deploy.

Which base model is recommended?

nvidia/parakeet-tdt-0.6b-v2 using stock NeMo SFT. The docs explicitly warn against fine-tuning the streaming Nemotron base (its SFT path collapses) and against the adapter-mixin path on TDT/RNNT decoders (produces NaN tensors).

How much improvement is realistic?

On the reference 39-row manifest the skill reports KER 0.513 to 0.128 in 3 epochs, a 75 percent relative reduction, with drug names improving the most. Production runs use 10 to 30 epochs with early stopping.

Is the cloud install script safe to run?

The reference flags the curl-piped install pattern as an antipattern and recommends downloading the script first, verifying it with shasum and inspecting it before running, pinning to a release tag, or using Homebrew on macOS.

All Files

7 files
evals/evals.json3.2 KB
View
references/stage4-finetune.md8.5 KB
View
SKILL.md19.5 KB
View
BENCHMARK.md3.7 KB
View
references/container-paths.md3.1 KB
View
skill-card.md3.7 KB
View
skill.oms.sig5.0 KB
View

Install digital-health-clinical-asr-finetune

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill

Repository

nvidia/skills