spark-engineer
Use when writing Spark jobs, debugging performance issues, or configuring cluster settings for Apache Spark applications, distributed data processing pipelines, or big data workloads. Invoke to write DataFrame transformations, optimize Spark SQL queries, implement RDD pipelines, tune shuffle operations, configure executor memory, process .parquet files, handle data partitioning, or build structured streaming analytics.
Security Assessment
About spark-engineer
The `spark-engineer` skill provides expert guidance and implementation support for building and optimizing Apache Spark applications. It is designed to help engineers handle large-scale distributed data processing pipelines efficiently, addressing common challenges such as performance bottlenecks, data skew, shuffle optimization, and cluster configuration. By leveraging this skill, users can implement high-performance Spark jobs, tune Spark SQL queries, manage RDD and DataFrame operations, and configure cluster resources to achieve scalable, production-ready pipelines.
Key features of this skill include writing and optimizing DataFrame transformations, implementing RDD pipelines, tuning shuffle operations, handling data partitioning, and building structured streaming analytics. It also provides practical guidance for configuring executor memory, managing persistence levels, broadcasting small dimension tables, and mitigating data skew using techniques like salting. Reference materials are organized by topic, including Spark SQL & DataFrames, RDD operations, partitioning and caching strategies, performance tuning, and streaming patterns, allowing users to load detailed guidance contextually.
This skill is primarily aimed at senior data engineers, Spark developers, and data platform specialists who work on big data workloads and distributed computing environments. It is particularly useful for teams managing ETL pipelines, analytics workflows, or real-time streaming applications. Use cases include building scalable batch and streaming pipelines, optimizing Spark jobs for large datasets, implementing production-grade Spark applications, and monitoring cluster performance to ensure resource efficiency and target performance levels.
FAQ
What types of Spark workloads is this skill best suited for?
It is suited for both batch and streaming workloads, including large-scale ETL pipelines, Spark SQL queries, RDD transformations, and structured streaming analytics.
Which Spark APIs and components are supported?
The skill supports the DataFrame API, Spark SQL, RDD operations, broadcast joins, caching, partitioning strategies, and structured streaming patterns.
Are there any requirements for cluster configuration or memory settings?
Yes, effective use of this skill involves tuning executor memory, shuffle partitions, and other Spark configurations to optimize performance and avoid data skew or shuffle spills.
Can this skill handle data skew and performance issues?
Yes, it provides strategies such as salting skewed keys, tuning shuffle partitions, and analyzing Spark UI metrics to detect and mitigate performance bottlenecks.
What are the limitations regarding input data formats?
The skill primarily references `.parquet` files and assumes structured data suitable for Spark SQL or DataFrame operations, though RDD transformations can support more generic formats.
Install spark-engineer
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
jeffallan/claude-skills