Back to Skills

harvard-art-museums-data-pipeline

Build end-to-end data engineering pipelines with the Harvard Art Museums API, ETL processes, SQL analytics, and Streamlit visualization

4stars1forksUpdated 7/29/2026

Security Assessment

Safe(92/100)
Security Score92/100

About harvard-art-museums-data-pipeline

Harvard Art Museums Data Pipeline is a data-engineering tutorial skill that shows how to build an end-to-end ETL and analytics application on top of the public Harvard Art Museums API. It solves the problem of turning a paginated, nested-JSON museum API into a normalized, queryable relational dataset with an interactive dashboard, walking through the full Extract → Transform → Load → Analyze → Visualize flow.

The skill documents cloning a companion GitHub project and installing Python dependencies (streamlit, pandas, requests, mysql-connector-python, plotly, python-dotenv). It covers configuration via a .env file holding a Harvard API key and MySQL/TiDB Cloud connection settings, database schema setup with normalized tables (artifact metadata, media, colors) and foreign keys, rate-limited paginated API collection, transformation of nested records into DataFrames for batch insertion, analytical SQL queries, and Plotly/Streamlit dashboards. Credential handling follows standard practice—secrets are read from environment variables loaded from a local .env file rather than hard-coded.

Target users are data-engineering learners, analysts, and developers who want a concrete, reproducible example of an API-to-warehouse-to-dashboard pipeline. It is authored as part of the "Data Skills" collection and is primarily instructional: much of its value is in the documented patterns and example code rather than a standalone automated tool, and the operations it describes (public API reads, database inserts, dashboard rendering) are routine and non-destructive.

FAQ

What does the pipeline actually build?

An ETL workflow that fetches artifact data from the Harvard Art Museums API, normalizes it into MySQL/TiDB tables, runs analytical SQL, and visualizes results in a Streamlit dashboard with Plotly charts.

What do I need to run it?

Python with the listed dependencies, a free Harvard Art Museums API key, and access to a MySQL-compatible database (local MySQL or TiDB Cloud). Credentials go in a local .env file.

How are API keys and database passwords handled?

They are stored in a .env file and read at runtime via python-dotenv/os.getenv, following standard environment-variable practice rather than being embedded in code.

Does it handle API pagination and rate limits?

Yes. The collection routine pages through results, stops when no further pages exist, and sleeps briefly between requests to respect rate limits.

Is this a finished tool or a learning template?

It is largely instructional—a documented reference implementation pointing to a companion GitHub repository—so expect to clone and adapt the example code rather than invoke a packaged tool.

Install harvard-art-museums-data-pipeline

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill