This fully revised edition teaches you to scale Python across cores, machines, and GPUs using the libraries you already know. I'll teach you to build Dask arrays, dataframes, bags and delayed graphs, and to tune partitions and chunks so the cluster actually earns its keep. I'll also show you to use the Dask expression system to let the optimiser push filters down and prune columns before a byte moves.
As you'll see, you'll be working with pandas 3.0 throughout, including copy-on-write, PyArrow-backed strings and the new expression API. There are dedicated chapters that take you through terabyte-scale multi-dimensional data in HDF5, NetCDF, TIFF and Zarr, with Xarray, Kerchunk and cloud-native datacubes. The good thing about Dask is that it's the lighter, and the quicker choice than Spark in most scenarios. After that, we'll move on to GPU acceleration with RAPIDS and cuDF, distributed XGBoost and dask-ml for scaled machine learning, and production deployment on Kubernetes with monitoring, security, and cost control. Every chapter builds on one ongoing project, so the skills you learn are added to over time instead of being spread out.
Key Learnings
Diagnose whether a workload needs parallelism, a bigger machine, or better code.
Size partitions and array chunks so clusters stop thrashing and start scaling.
Read task graphs and dashboard panels to locate bottlenecks within minutes.
Exploit the Dask expression system for predicate pushdown and column projection.
Migrate pandas code to 3.0 copy-on-write and PyArrow-backed string dtypes.
Stream terabyte HDF5, NetCDF, TIFF, and Zarr archives without exhausting memory.
Build cloud-native, versioned datacubes using Zarr v3, Kerchunk, and Icechunk.
Choose between Dask, Spark, Ray, and Polars using measured, honest criteria.
Accelerate dataframes and gradient boosting on GPUs with RAPIDS and dask-cuda.
Deploy, secure, monitor, and cost-control Dask clusters on production Kubernetes.
Table of Content
Up and Running with Dask
Working with Dask Collections
Tuning Partitions and Chunks
Optimizing Queries with Expression System
Scaling Out with Distributed Clusters
Scaling pandas 3.0 with Dask
Reading and Writing Data at Scale
Working with Labeled Arrays in Xarray
Reading HDF5 and NetCDF Archives
Scaling Imagery with TIFF
Building Cloud-native Datacubes with Zarr
Accelerating Dask with GPUs
Operating Dask in Production
Target Audience
If you're wondering whether your Pandas script is stalling and you've tried adding memory but only got a few weeks out of it, then this latest edition on running parallel operations should help you figure out what to do next. You just need to have a basic understanding of Python programming. That's all you need to get reading this book!