Programming Concurrency on GPUs with Rust - Zyle Kot

Programming Concurrency on GPUs with Rust

By: Zyle Kot

eBook | 12 September 2026

At a Glance

eBook


$42.99

or 4 interest-free payments of $10.75 with

 or 

Instant Digital Delivery to your Kobo Reader App

The bottleneck is not your kernel. In fact, it almost never is.

After spending enough time in Nsight, the way you read a timeline becomes as intuitive as a mechanic listening to an engine. The same faults keep cropping up, such as a GPU with thousands of cores sitting idle behind a blocking copy, a single default stream or a host thread frozen at a synchronisation point that was never necessary. This book is about identifying that fault early and eliminating it.

Designed for programming professionals, this book is written for GPU developers, Rust engineers and AI practitioners who want their devices to be busy, not just correct. Our approach is to stay in pure Rust throughout, using cuda-oxide for SIMT kernels and cutile-rs for the tile model. We also rely heavily on the type system to ensure that data races fail at compile time rather than at 3 AM during production. We will overlap transfers with compute using streams and pinned memory, build a staging ring that pipelines every stage and feed the GPU through Tokio producers with bounded-channel backpressure. Each technique is profiled, verified against a CPU reference and measured.

Key Learnings

Profile before you optimise, and let Nsight Systems show you who is actually waiting.

Write SIMT kernels in pure Rust where data races fail at compile time.

Treat default stream as the enemy of overlap, and fork your own.

Overlap copies with compute using pinned memory and a staging ring.

Reach for reduction ladder instead of atomics when summing on the device.

Express element-wise work as tile kernels, and let compiler map the hardware.

Keep the GPU fed with Tokio producers and bounded-channel backpressure.

Verify every kernel against CPU reference before you trust its timing.

Weigh throughput against latency, since tuning for one quietly taxes the other.

Know when to stop, because past a point concurrency only adds bookkeeping.

Table of Content

Rust's Threads, Timing and Ownership

CPU versus GPU Execution

SIMT Kernels, Blocks and Grids

Race-Free Parallel Kernels

Asynchronous GPU Execution

Overlapping Transfers with Streams

Tile Programming with cutile-rs

Async GPU Pipelines with Tokio

Throughput, Latency and Bottlenecks

on

More in Computer Programming & Software Development

The End of Leadership - Barbara Kellerman

eBOOK

This Is a Title - Josh Brody

eBOOK

$14.99

Trusted Intelligence - John S Pritchett

eBOOK

RRP $17.59

$16.99