harvard-art-museums-data-engineering-app
End-to-end data engineering and analytics application for Harvard Art Museums API with ETL pipelines, SQL analytics, and Streamlit visualization
Security Assessment
About harvard-art-museums-data-engineering-app
The Harvard Art Museums Data Engineering App is a catalogue-style skill that walks through building an end-to-end data engineering and analytics application on top of the Harvard Art Museums API. It demonstrates a complete API-to-visualization pipeline and points to an upstream reference implementation on GitHub, so it functions largely as a documented, copyable project template rather than a self-contained tool with its own runtime.
The documented pipeline follows API to ETL to SQL to analytics to visualization: it collects artifact data from the Harvard Art Museums API with pagination and rate limiting, transforms nested JSON into normalized relational tables (artifact metadata, media, and colors), loads the data into MySQL or TiDB Cloud databases, runs 20+ predefined SQL analytical queries, and visualizes the results through interactive Plotly dashboards in Streamlit. The SKILL.md includes concrete setup steps — cloning the upstream repo, installing dependencies (streamlit, pandas, requests, mysql-connector-python, plotly, python-dotenv), and configuring credentials through environment variables (API key and database host/user/password/name) loaded via python-dotenv — plus a normalized SQL schema with foreign keys and worked Python code for the extract, transform, and load stages.
It targets data engineering learners, students, and practitioners who want a realistic, reproducible example project covering API ingestion, relational modeling, SQL analytics, and dashboarding. It is best suited to those building a portfolio project or learning the ETL-to-dashboard workflow rather than a plug-and-play production system.
FAQ
Is this a standalone skill or a pointer to a project?
It is largely a documented walkthrough that references an upstream GitHub repository (Manali0711/Harvard-Artifacts-Collection-Data-Engineering-Analytics-App). It provides setup steps, schema, and ETL code samples rather than a fully bundled runtime.
What technologies does it use?
Python with streamlit, pandas, requests, mysql-connector-python, plotly, and python-dotenv, backed by a MySQL or TiDB Cloud database, and the public Harvard Art Museums API.
How are credentials handled?
Through environment variables loaded with python-dotenv: a Harvard API key (free, from the museum's API page) plus database host, user, password, and name. Credentials are read from the environment, not hard-coded.
What can I build with it?
An end-to-end pipeline that ingests artifact data with pagination and rate limiting, normalizes nested JSON into three related tables, loads it into SQL, runs 20+ analytical queries, and renders interactive Plotly dashboards in Streamlit.
Install harvard-art-museums-data-engineering-app
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
aradotso/data-skills