Get Free Shipping on orders over $79
GPU Programming using Rust and CUDA : Exploring Rust's potential in GPU and parallel computing using Rust-CUDA, cuda-oxide, and RustaCUDA - Maris Fenlor

GPU Programming using Rust and CUDA

Exploring Rust's potential in GPU and parallel computing using Rust-CUDA, cuda-oxide, and RustaCUDA

By: Maris Fenlor

Paperback | 25 July 2026

At a Glance

Paperback


RRP $126.49

$124.75

or 4 interest-free payments of $31.19 with

 or 

Ships in 5 to 10 business days

C++ has been the go-to for GPU programming for almost 20 years. Can Rust do the job, and how well?

This book is all about getting hands-on with different toolchains that connect Rust to NVIDIA hardware. There's RustaCUDA for safe host-side control, the Rust-CUDA project for writing kernels in pure Rust, and NVIDIA's experimental cuda-oxide compiler with its typed launches and async execution graphs.

We're going to build one Cargo workspace that keeps on growing. It'll include device queries, launch planning, Rust-written kernels, memory optimization, parallel reductions and scans, multi-stream pipelines, matrix multiplication benchmarked against cuBLAS, a Monte Carlo option pricer validated against a closed formula, and a complete batched inference application measured against a Python baseline. We'll check every result against a CPU reference, and the reports will give accurate numbers, including where libraries outperform hand-written kernels and where experimental toolchains are still a work in progress.

Key Learnings

Launch, synchronize, and verify GPU kernels with ownership-managed device memory.

Write real CUDA kernels using Rust-CUDA and cuda-oxide.

Plan grids, blocks, and warps for 2D workloads.

Accelerate transfer speeds with pinned memory and coalesced access patterns.

Build race-free thread cooperation using shared memory, barriers, and atomics.

Overlap transfers with computation using streams, events, and async Rust pipelines.

Optimize matrix multiplication and benchmark against cuBLAS ceiling.

Wrap CUDA C library safely with handles, error enums, and Drop.

Ship complete batched GPU inference application against Python baselines.

Diagnose performance with Nsight Systems, Nsight Compute, and compute-sanitizer.

Table of Content

New Beneficiary of GPU Computing

Thinking in Threads

Commanding GPU

Writing GPU Kernels

Cleaner Kernels with cuda-oxide

Mastering GPU Memory

Making Threads Cooperate

Keeping GPU Busy

Delivering Real Math

Borrowing NVIDIA's Muscle

Shipping Complete GPU Application

Proving Performance

More in Algorithms & Data Structures

Addiction by Design : Machine Gambling in Las Vegas - Natasha Dow Schull
Learning Algorithms : A Programmer's Guide to Writing Better Code - George Heineman
Python for Algorithmic Trading : From Idea to Cloud Deployment - Yves Hilpisch
Code Dependent : Living in the Shadow of AI - Madhumita Murgia

RRP $24.99

$21.75

13%
OFF
How to Prove It : A Structured Approach - Daniel J. Velleman

RRP $73.95

$70.75

Finite Element Mesh Generation, 2e - Daniel S.H.  Lo
Fundamentals of Data Structures and Algorithms - Elvis C. Foster

RRP $158.00

$121.99

23%
OFF
Hacker's Delight - Henry Warren

RRP $97.60

$76.75

21%
OFF
Python Using GPT-5 and Gemini - Oswald Campesato