Get Free Shipping on orders over $0
Apache Spark for Distributed Systems Engineers - Devlin Nexley

Apache Spark for Distributed Systems Engineers

By: Devlin Nexley

Paperback | 24 September 2026

At a Glance

Paperback


$31.89

or 4 interest-free payments of $7.97 with

 or 

Ships in 5 to 7 business days

Apache Spark is not just a data processing framework-it is a distributed execution system built on deep principles of cluster computing, DAG-based scheduling, memory-aware computation, and fault-tolerant design.

This book provides a rigorous, systems-level examination of Spark as a production-grade distributed execution engine. It moves beyond APIs, tutorials, and surface-level usage patterns to expose the internal mechanisms that govern how Spark actually executes workloads at scale.

Designed for experienced engineers working in distributed systems, backend infrastructure, and data platform engineering, this book dissects Spark as an architectural system rather than a development tool.

Inside, you will explore how Spark transforms high-level computations into distributed execution graphs, how it schedules and coordinates work across clusters, and how it manages the complexity of large-scale data movement in cloud-native environments.

Key areas covered include:

Internal architecture of Spark's driver, executors, and cluster coordination model

DAG construction, stage decomposition, and task scheduling mechanics

Shuffle architecture, data movement patterns, and network bottlenecks

Memory management, execution optimization, and JVM runtime behavior

Fault tolerance through lineage reconstruction and retry semantics

Query execution via Spark SQL, Catalyst optimizer, and Tungsten engine

Structured Streaming and incremental computation models

Performance bottlenecks, skew handling, and production tuning strategies

Cloud-native execution on object storage systems and Kubernetes

Integration with modern lakehouse ecosystems such as Delta Lake, Iceberg, and Hudi

Rather than presenting Spark as a tool to be used, this book treats it as a distributed systems case study-revealing how large-scale data infrastructure is engineered, optimized, and operated under real production constraints.

By the end, readers will understand not only how Spark works, but why its architecture is designed the way it is, what trade-offs shape its execution model, and how it fits into the broader evolution of modern distributed data platforms.

This is a book for engineers who build systems, not just pipelines.

More in Databases

Statistics and Data Handling for Biologists : A Student's Guide - Neil Millar
Dynamics of Marine Structures - Yingguang  Wang

RRP $503.95

$442.99

12%
OFF
Data Empire : How information shaped human history - Roopika Risam
Social Research Methods : 4th Edition - Maggie Walter

RRP $101.95

$87.75

14%
OFF
Data Analytics for Accounting ISE : 3rd Edition - Vernon J. Richardson

RRP $169.95

$146.75

14%
OFF
Vector Databases : A Practical Introduction - Nitin Borwankar

RRP $133.00

$64.75

51%
OFF
Data-driven BIM for Energy Efficient Building Design : 1st Edition - Saeed Banihashemi
Building a Scalable Data Warehouse with Data Vault 2.0 - Dan Linstedt
Microsoft Power BI For Dummies : For Dummies (Computer/Tech) - Jack A. Hyman