tao-finetune-cosmos-reason
Cosmos3-Nano video QA supervised fine-tuning with FSDP parallelism. Use when training or evaluating video question-answering models, fine-tuning Cosmos3-Nano or compatible Cosmos Reason models with SFT/LoRA, or working with Cosmos-RL. Trigger phrases include "fine-tune Cosmos", "Cosmos3 Nano Reasoner", "Cosmos-RL SFT", "video QA fine-tune", "Cosmos3-Nano training".
Security Assessment
About tao-finetune-cosmos-reason
This NVIDIA skill drives supervised fine-tuning (SFT) and evaluation of Cosmos Reason video question-answering models, with the packaged default base model nvidia/Cosmos3-Nano. It solves the complexity of configuring a multi-GPU video-QA training run: it standardizes how the base model is sourced from HuggingFace, how FSDP parallelism is expressed (dp_shard_size for GPU count and dp_replicate_size for node count instead of the usual num_gpus/num_nodes), and when a checkpoint conversion to a Qwen3-VL HF safetensors directory is needed because some Cosmos-RL images cannot load the native Cosmos3 Omni format directly.
The skill is schema-driven and AutoML-aware. Generated TAO Core dataclass schemas live in schemas/<action>.schema.json with a manifest, and each emits a spec_template_<action>.yaml. Reference guides cover launch intake and preflight, evaluation (task types, LoRA eval, selective download, results), and AutoML/HPO policy including mapping objectives like accuracy targets to evaluation metrics and search spaces over learning rate, batch size, and weight decay. Train requests default to AutoML-on and route through a shared tao-run-automl runner when schemas and templates are packaged, falling back to direct training when AutoML is disabled or schemas are missing. Non-train actions (evaluate, inference, quantize) stay in this skill. It requires Docker with the NVIDIA container toolkit, and credentials are handled through the standard HF_TOKEN environment variable, with the user directed to accept the model license and provide a read token for gated models.
Target users are ML engineers and researchers fine-tuning or evaluating Cosmos Reason / Cosmos3-Nano video-QA models. Use cases include SFT and LoRA fine-tuning, accuracy-driven hyperparameter search, and evaluation of video-QA reasoning models.
FAQ
What does this skill fine-tune?
Cosmos Reason video question-answering models, with the packaged default base model hf_model://nvidia/Cosmos3-Nano, using supervised fine-tuning (SFT) and LoRA with FSDP-based parallelism.
What are the system requirements?
Per its metadata it requires Docker plus the NVIDIA container toolkit. Pretrained weights are pulled from HuggingFace, and gated models require an HF_TOKEN with read access after accepting the model agreement.
How is GPU and node count specified?
It uses FSDP parallelism: dp_shard_size sets the GPU count and dp_replicate_size sets the node count, rather than the standard num_gpus/num_nodes fields.
Do I need to convert the base model?
Only for certain Cosmos-RL images that cannot load the native Cosmos3 Omni checkpoint format. In that case you convert Cosmos3-Nano to a Qwen3-VL HF safetensors directory and use that as the pretrained-model path.
How does AutoML work here?
The model is AutoML-enabled by default; train requests route through the shared tao-run-automl runner when the train schema and template are packaged. You can disable it per run (automl_policy: off) for plain training.
All Files
25 filesInstall tao-finetune-cosmos-reason
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
nvidia/skills