The bottleneck is not your kernel. In fact, it almost never is.
After spending enough time in Nsight, the way you read a timeline becomes as intuitive as a mechanic listening to an engine. The same faults keep cropping up, such as a GPU with thousands of cores sitting idle behind a blocking copy, a single default stream or a host thread frozen at a synchronisation point that was never necessary. This book is about identifying that fault early and eliminating it.
Designed for programming professionals, this book is written for GPU developers, Rust engineers and AI practitioners who want their devices to be busy, not just correct. Our approach is to stay in pure Rust throughout, using cuda-oxide for SIMT kernels and cutile-rs for the tile model. We also rely heavily on the type system to ensure that data races fail at compile time rather than at 3 AM during production. We will overlap transfers with compute using streams and pinned memory, build a staging ring that pipelines every stage and feed the GPU through Tokio producers with bounded-channel backpressure. Each technique is profiled, verified against a CPU reference and measured.
Key Learnings
Profile before you optimise, and let Nsight Systems show you who is actually waiting.
Write SIMT kernels in pure Rust where data races fail at compile time.
Treat default stream as the enemy of overlap, and fork your own.
Overlap copies with compute using pinned memory and a staging ring.
Reach for reduction ladder instead of atomics when summing on the device.
Express element-wise work as tile kernels, and let compiler map the hardware.
Keep the GPU fed with Tokio producers and bounded-channel backpressure.
Verify every kernel against CPU reference before you trust its timing.
Weigh throughput against latency, since tuning for one quietly taxes the other.
Know when to stop, because past a point concurrency only adds bookkeeping.
Table of Content
Rust's Threads, Timing and Ownership
CPU versus GPU Execution
SIMT Kernels, Blocks and Grids
Race-Free Parallel Kernels
Asynchronous GPU Execution
Overlapping Transfers with Streams
Tile Programming with cutile-rs
Async GPU Pipelines with Tokio
Throughput, Latency and Bottlenecks