Parallel Python with Dask, Second Edition : Scale pandas, NumPy, and Xarray across cores, clusters, and GPUs for terabyte-scale analytics and machine learning - Patrick J

Parallel Python with Dask, Second Edition

Scale pandas, NumPy, and Xarray across cores, clusters, and GPUs for terabyte-scale analytics and machine learning

By: Patrick J

eBook | 5 September 2026

At a Glance

eBook


RRP $63.79

$61.99

or 4 interest-free payments of $15.50 with

 or 

Instant Digital Delivery to your Kobo Reader App

This fully revised edition teaches you to scale Python across cores, machines, and GPUs using the libraries you already know. I'll teach you to build Dask arrays, dataframes, bags and delayed graphs, and to tune partitions and chunks so the cluster actually earns its keep. I'll also show you to use the Dask expression system to let the optimiser push filters down and prune columns before a byte moves.

As you'll see, you'll be working with pandas 3.0 throughout, including copy-on-write, PyArrow-backed strings and the new expression API. There are dedicated chapters that take you through terabyte-scale multi-dimensional data in HDF5, NetCDF, TIFF and Zarr, with Xarray, Kerchunk and cloud-native datacubes. The good thing about Dask is that it's the lighter, and the quicker choice than Spark in most scenarios. After that, we'll move on to GPU acceleration with RAPIDS and cuDF, distributed XGBoost and dask-ml for scaled machine learning, and production deployment on Kubernetes with monitoring, security, and cost control. Every chapter builds on one ongoing project, so the skills you learn are added to over time instead of being spread out.

Key Learnings

Diagnose whether a workload needs parallelism, a bigger machine, or better code.

Size partitions and array chunks so clusters stop thrashing and start scaling.

Read task graphs and dashboard panels to locate bottlenecks within minutes.

Exploit the Dask expression system for predicate pushdown and column projection.

Migrate pandas code to 3.0 copy-on-write and PyArrow-backed string dtypes.

Stream terabyte HDF5, NetCDF, TIFF, and Zarr archives without exhausting memory.

Build cloud-native, versioned datacubes using Zarr v3, Kerchunk, and Icechunk.

Choose between Dask, Spark, Ray, and Polars using measured, honest criteria.

Accelerate dataframes and gradient boosting on GPUs with RAPIDS and dask-cuda.

Deploy, secure, monitor, and cost-control Dask clusters on production Kubernetes.

Table of Content

Up and Running with Dask

Working with Dask Collections

Tuning Partitions and Chunks

Optimizing Queries with Expression System

Scaling Out with Distributed Clusters

Scaling pandas 3.0 with Dask

Reading and Writing Data at Scale

Working with Labeled Arrays in Xarray

Reading HDF5 and NetCDF Archives

Scaling Imagery with TIFF

Building Cloud-native Datacubes with Zarr

Accelerating Dask with GPUs

Operating Dask in Production

Target Audience

If you're wondering whether your Pandas script is stalling and you've tried adding memory but only got a few weeks out of it, then this latest edition on running parallel operations should help you figure out what to do next. You just need to have a basic understanding of Python programming. That's all you need to get reading this book!

on

More in Software Engineering

The End of Leadership - Barbara Kellerman

eBOOK

This Is a Title - Josh Brody

eBOOK

$14.99

Trusted Intelligence - John S Pritchett

eBOOK

RRP $17.59

$16.99