nemo-data-designer-plugin
Use when the user wants to create a dataset, generate synthetic data, or build a data generation pipeline.
Security Assessment
About nemo-data-designer-plugin
nemo-data-designer-plugin is an NVIDIA (NeMo Platform) skill for creating synthetic datasets and data-generation pipelines using the Data Designer library. It activates when a user wants to create a dataset, generate synthetic data, or build a generation pipeline, and it drives an end-to-end workflow that produces a runnable Python script defining a load_config_builder() function that returns a DataDesignerConfigBuilder. The problem it solves is turning a natural-language dataset description into a validated, previewable configuration without the user needing to memorize Data Designer's config schema.
The skill offers two modes: Autopilot (make reasonable decisions autonomously when the user signals they don't want questions) and Interactive (the default, driven by a separate workflow file). The documented workflow resolves the nemo CLI, learns available context via nemo data-designer agent context, infers diversity axes and schema, plans columns/samplers/processors/validators, writes the script with PEP 723 inline dependency metadata, then validates, previews with saved sample records (surfaced as clickable HTML), and optionally creates a full dataset of N records. Reference files cover person sampling, seed datasets, preview review, and NeMo Platform plugin additions such as sourcing model configs from IGW providers. The skill is explicit about safety and permission: it will not install the CLI or dependencies without the user's permission, and its troubleshooting guidance asks before retrying commands with a sandbox disabled. Its skill card notes no external credential is required and advises against putting secrets in prompts or logs.
Target users are developers and ML engineers who need synthetic training data, evaluation datasets, or reproducible data pipelines, especially those working within the NVIDIA NeMo Platform ecosystem. It requires Python 3.11+ and the nemo data-designer CLI.
FAQ
What does this skill produce?
A descriptively named Python file in the current directory with a load_config_builder() function returning a DataDesignerConfigBuilder, using PEP 723 inline metadata for dependencies.
What is the difference between Autopilot and Interactive modes?
Autopilot makes reasonable design decisions autonomously (used when the user says things like "you decide" or "just build it"); Interactive is the default and asks clarifying questions. Only the matching workflow file is read.
What are the prerequisites?
The nemo data-designer CLI, which requires Python 3.11 or newer. If the CLI is not found the skill tells the user and asks permission before installing anything.
Does it need API keys or credentials?
The skill card lists no required external credential. It advises not putting secrets in prompts, logs, or output and using least-privilege credentials. Model configs can be sourced from IGW providers per the plugin additions reference.
How does it verify data quality?
It validates the script, then runs a preview that saves sample records as HTML for review, checking for issues like mode collapse, instruction compliance, and encoding integrity before optionally generating the full dataset.
All Files
12 filesInstall nemo-data-designer-plugin
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
nvidia/skills