harvard-artifacts-collection-etl-analytics
Build ETL pipelines and analytics dashboards for Harvard Art Museums API data using Python, SQL, and Streamlit
Security Assessment
About harvard-artifacts-collection-etl-analytics
harvard-artifacts-collection-etl-analytics is a hands-on skill for building ETL pipelines and analytics dashboards from the Harvard Art Museums API using Python, SQL, and Streamlit. It closely parallels the companion Harvard artifacts skill, addressing the same goal of turning a public museum API into a queryable database and interactive dashboard, and describes the architecture as API to ETL to SQL to Analytics to Visualization.
The skill supplies runnable Python: a fetch function that pages the API (filtering to artifacts with images) with rate limiting, extraction functions that split artifacts into metadata, media, and color DataFrames, and a loading layer that creates a MySQL/TiDB schema via mysql-connector-python. It documents obtaining a free Harvard API key and storing configuration in environment variables (API key and database host, user, password, name) exported in the shell, which is standard credential-setup guidance. It highlights 20+ analytical queries and auto-generated visualizations with Plotly.
It targets data engineers, analysts, and learners who want a cultural-heritage-data reference implementation for pagination, JSON-to-relational transformation, schema design, and dashboarding. Common use cases include practicing ETL construction, modeling nested API responses into normalized tables, and building a Streamlit analytics front end. Like its sibling, it directs users to clone an upstream GitHub repository for the complete application, so it operates partly as a guided entry point to that external project.
FAQ
How does this differ from the sibling Harvard skill?
It covers the same Harvard Art Museums ETL-and-analytics workflow with slightly different extraction fields and setup emphasis; both point to the same upstream application repository.
How is the API key stored?
In an environment variable (HARVARD_API_KEY), exported in the shell alongside database connection variables, which is standard credential setup with no data exfiltration.
What is the technology stack?
Python (requests, pandas) for ETL, MySQL or TiDB Cloud via mysql-connector-python for storage, and Streamlit plus Plotly for visualization.
Does the API access respect rate limits?
Yes. The fetch/collect functions add a delay between page requests to stay within API rate limits.
Is it self-contained?
It embeds core code but instructs cloning an upstream GitHub project for the full app, so it functions in part as a pointer to that repository.
Install harvard-artifacts-collection-etl-analytics
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
aradotso/data-skills