Back to Skills

data-context-extractor

Generate or improve a company-specific data analysis skill by extracting tribal knowledge from analysts. BOOTSTRAP MODE - Triggers: "Create a data context skill", "Set up data analysis for our warehouse", "Help me create a skill for our database", "Generate a data skill for [company]" → Discovers schemas, asks key questions, generates initial skill with reference files ITERATION MODE - Triggers: "Add context about [domain]", "The skill needs more info about [topic]", "Update the data skill with

21,707stars2,536forksUpdated 6/22/2026

Security Assessment

Safe(100/100)
Security Score100/100

About data-context-extractor

The data-context-extractor is a meta-skill that helps data analysts teach Claude about their company's specific data warehouse, terminology, and metrics. Rather than using generic SQL knowledge, it extracts tribal knowledge from analysts through guided conversations and generates a tailored data analysis skill that understands your organization's unique data landscape. The result is a reusable skill that any analyst on your team can use to get contextually accurate answers about your data.

The skill operates in two distinct modes. Bootstrap Mode creates a new data analysis skill from scratch by connecting to your data warehouse, exploring schemas, and asking targeted questions about entity definitions, primary identifiers, key metrics, data hygiene filters, and common gotchas. Iteration Mode loads an existing skill and adds or updates domain-specific reference files, allowing you to progressively enrich the skill over time without starting over. Both modes produce structured reference files covering entities, metrics, and per-domain table documentation.

This skill is designed for data analysts, analytics engineers, and data teams who want Claude to reason accurately about their company's data without having to re-explain context in every conversation. It supports major data warehouses including BigQuery, Snowflake, PostgreSQL, Redshift, and Databricks, and integrates with available MCP tools in the session to query schemas directly.

FAQ

What data warehouses are supported?

BigQuery, Snowflake, PostgreSQL, Redshift, and Databricks are explicitly supported. The skill uses available MCP tools in the current session to connect, so any warehouse with a compatible MCP tool should work.

What is the difference between Bootstrap Mode and Iteration Mode?

Bootstrap Mode creates a brand-new data analysis skill from scratch, including schema discovery and initial reference files. Iteration Mode loads an existing skill and appends or updates specific reference files with new domain knowledge, so you can improve the skill incrementally.

What does the generated skill output look like?

The output is a directory named [company]-data-analyst/ containing a SKILL.md and a references/ folder with files for entities, metrics, and per-domain table documentation.

Do I need to answer all questions at once?

No. The skill asks questions conversationally across several phases, not all at once. You can provide as much or as little detail as you have available, and use Iteration Mode later to fill in gaps.

Can multiple analysts contribute to the same skill?

Yes. Iteration Mode is designed for this — different analysts can add context about their domains over time, progressively improving the shared skill.

All Files

6 files
references/sql-dialects.md4.5 KB
View
references/domain-template.md3.4 KB
View
references/skill-template.md3.7 KB
View
SKILL.md7.0 KB
View
references/example-output.md6.1 KB
View
scripts/package_data_skill.py3.7 KB
View

Install data-context-extractor

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill