Back to Skills

harvard-artifacts-collection-data-engineering-analytics

End-to-end data engineering and analytics application using Harvard Art Museums API with ETL pipelines, SQL analytics, and Streamlit visualization

4stars1forksUpdated 7/31/2026

Security Assessment

Safe(92/100)
Security Score92/100

About harvard-artifacts-collection-data-engineering-analytics

harvard-artifacts-collection-data-engineering-analytics is a tutorial-style skill that walks through building an end-to-end data engineering and analytics application on top of the Harvard Art Museums API. It solves the problem of demonstrating a complete ETL-to-dashboard workflow by extracting artifact data from the API, transforming nested JSON into structured relational tables, loading it into a SQL database (MySQL or TiDB Cloud), and visualizing it with Streamlit and Plotly.

The skill provides concrete Python code: an extract step that paginates the API with rate limiting, a transform step that flattens artifacts into metadata, media, and color DataFrames, and a load step that creates a normalized schema (artifactmetadata, artifactmedia, artifactcolors with foreign keys) via mysql-connector-python. Configuration uses a .env file for the Harvard API key and database host, port, user, password, and name, following standard credential-setup practice with python-dotenv. It advertises 20+ analytical SQL queries and interactive visualizations, and points to an upstream companion repository for the full application.

It targets data engineering learners, analytics developers, and anyone wanting a reference implementation for API-driven ETL and dashboarding. Typical use cases include learning ETL pipeline design, practicing relational schema modeling from nested JSON, and standing up a Streamlit analytics dashboard. Because the skill largely mirrors and links to an external GitHub project, it functions partly as a guided catalogue entry into that upstream repo rather than a fully self-contained tool.

FAQ

What is the data source?

The Harvard Art Museums API, accessed with a free API key and page-based requests that include rate limiting.

What database and tools does it use?

MySQL or TiDB Cloud via mysql-connector-python for storage, with Python (requests, pandas) for ETL and Streamlit plus Plotly for the dashboard.

How are credentials handled?

Through a .env file (Harvard API key and database connection settings) loaded with python-dotenv, which is standard local credential setup with no exfiltration.

Is the full application included in the skill?

The skill embeds representative code but instructs cloning an upstream companion GitHub repository for the complete application, so it partly serves as a pointer to that project.

What can I build with it?

A normalized artifact database (metadata, media, colors), 20+ analytical SQL queries, and interactive Streamlit/Plotly visualizations of museum collection data.

Install harvard-artifacts-collection-data-engineering-analytics

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill