Get Free Shipping on orders over $79
NVIDIA GPU Infrastructure Fundamentals : A structured guide to NVIDIA GPU infrastructure, from CUDA to production operations - Vivian Aranha

NVIDIA GPU Infrastructure Fundamentals

A structured guide to NVIDIA GPU infrastructure, from CUDA to production operations

By: Vivian Aranha

eText | 31 August 2026 | Edition Number 1

At a Glance

eText


$54.99

or 4 interest-free payments of $13.75 with

 or 

Instant online reading in your Booktopia eTextbook Library *

Why choose an eTextbook?

Instant Access *

Purchase and read your book immediately

Read Aloud

Listen and follow along as Bookshelf reads to you

Study Tools

Built-in study tools like highlights and more

* eTextbooks are not downloadable to your eReader or an app and can be accessed via web browsers only. You must be connected to the internet and have no technical issues with your device or browser that could prevent the eTextbook from operating.

Decode the NVIDIA GPU ecosystem in one structured guide. Compare technologies, understand how platform layers interact, and build the judgment to evaluate infrastructure choices and trade-offs.

Key Features

  • Understand how GPUs, CUDA, networking, storage, and DPUs support AI workloads
  • Learn the roles of MIG, vGPU, DCGM, Kubernetes, Slurm, NGC, and Triton
  • Connect infrastructure components across the AI development and deployment lifecycle

Book Description

NVIDIA GPU infrastructure spans hardware, system software, networking, storage, orchestration, MLOps, and inference. Understanding how these components fit together, where their responsibilities overlap, and which distinctions matter requires a clear, structured path. This book provides that path through one coherent narrative of the NVIDIA GPU infrastructure stack. It covers accelerated computing, CUDA, and the NVIDIA software ecosystem before comparing data center GPUs against workload characteristics. You will examine MIG, vGPU, and DCGM for resource sharing and monitoring; Kubernetes and Slurm for GPU scheduling; and the networking and storage layer, including Ethernet, InfiniBand, RDMA, GPUDirect Storage, BlueField DPUs, and DOCA. Later chapters connect infrastructure to the AI lifecycle through Airflow, MLflow, and Kubeflow for MLOps, NGC for software delivery, and ONNX, TensorRT, and Triton for inference. You will also explore the Kubernetes components, monitoring technologies, scaling considerations, and diagnostic concepts that support production GPU clusters. By connecting these technologies instead of presenting them as isolated products, the book helps you compare platform choices, understand component boundaries, discuss trade-offs, and develop a durable mental model of NVIDIA GPU infrastructure.

What you will learn

  • Distinguish AI, machine learning, and deep learning
  • Explain why GPUs accelerate modern AI workloads
  • Match NVIDIA GPUs to training and inference requirements
  • Select MIG or vGPU for common resource-sharing scenarios
  • Compare Ethernet and InfiniBand for distributed AI workloads
  • Map MLOps tools to the right stage of the AI lifecycle
  • Differentiate ONNX, TensorRT, and Triton in inference workflows
  • Trace GPU cluster issues across platform layers

Who this book is for

This book is for system administrators, cloud and DevOps professionals, data center and networking teams, solution architects, technical managers, presales professionals, and beginners who need a clear understanding of NVIDIA GPU infrastructure. It is especially relevant to professionals moving into AI infrastructure roles, evaluating GPU platform technologies, or collaborating across compute, networking, MLOps, and operations teams. Basic familiarity with IT, cloud, or data center concepts is helpful; programming, data science, and previous GPU experience are not required.

on
Desktop
Tablet
Mobile

More in Parallel Processing