enterprise-data-engineering-pipeline-ssis-pyspark
End-to-end ELT pipeline using SSIS, SQL Server, and PySpark for enterprise data warehousing and analytics
Security Assessment
About enterprise-data-engineering-pipeline-ssis-pyspark
Enterprise Data Engineering Pipeline (SSIS + PySpark) is a reference skill documenting an end-to-end ELT solution for enterprise data warehousing and analytics. It solves the problem of assembling a full pipeline from disparate Microsoft and big-data tools by combining SSIS for ETL orchestration, SQL Server with a star-schema data warehouse, Python and Pandas for data-quality audits and visualization, and PySpark for large-scale analytics and aggregation. The pipeline ingests raw CSV files (Sales, Products, Customers), transforms them through SSIS, loads them into a dimensional model, and performs analytics at scale.
The documentation lays out the architecture in layers — a source layer of raw CSVs, an ETL layer of SSIS packages handling extraction, transformation, and error handling, a SQL Server storage layer with fact and dimension tables, and an analytics layer of Python and PySpark scripts. It provides prerequisites and setup (SQL Server 2019+, SSIS, SSDT, Python 3.10+, Java 8+ for PySpark), the pip dependencies, and SQL DDL that creates the warehouse database, dimension and fact tables, and business-intelligence views such as revenue by product and customer lifetime value.
It targets enterprise data engineers and BI developers working in the Microsoft data stack who need a worked blueprint for a star-schema warehouse fed by SSIS and analyzed with PySpark. As a documentation-and-reference skill describing standard warehouse setup, DDL, and analytics scripts with no destructive or exfiltration behavior, it is benign educational tooling.
FAQ
What does the pipeline do?
It ingests raw CSV files for Sales, Products, and Customers, transforms them via SSIS, loads them into a SQL Server star-schema warehouse, and runs analytics with Python/Pandas and PySpark.
What are the prerequisites?
SQL Server 2019+ (Developer or Enterprise), SQL Server Integration Services, Visual Studio with SSDT, Python 3.10+, and Java 8+ for PySpark, plus Python packages including pandas, sqlalchemy, pyodbc, pyspark, and matplotlib.
What warehouse design does it use?
A star schema with dimension tables (Customers, Products) and a fact table (Sales), along with BI views such as revenue by product and customer lifetime value.
Why combine SSIS and PySpark?
SSIS handles ETL orchestration and error handling into SQL Server, while PySpark provides big-data analytics and aggregation that scales to millions of rows.
Is this a runnable project or a reference?
It is a reference blueprint with setup instructions, SQL DDL, and script structure demonstrating the enterprise pipeline pattern.
Install enterprise-data-engineering-pipeline-ssis-pyspark
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
aradotso/data-skills