Get Free Shipping on orders over $0
vLLM Serving : Highâ'Throughput LLM APIs with PagedAttention and KV Cache Tuning - Trex Team

vLLM Serving

Highâ'Throughput LLM APIs with PagedAttention and KV Cache Tuning

By: Trex Team

Paperback | 24 June 2026

At a Glance

Paperback


$76.75

or 4 interest-free payments of $19.19 with

 or 

Ships in 5 to 10 business days

"vLLM Serving: Highâ'Throughput LLM APIs with PagedAttention and KV Cache Tuning"

Built for experienced ML systems engineers, platform architects, and performance-minded practitioners, this book is a deep technical guide to serving large language models with vLLM at production scale. Rather than treating inference as a black box, it explains the real control surfaces behind throughput, latency, and memory efficiency. Readers who already know LLM fundamentals but want to reason rigorously about serving behavior will find an internals-first, systems-oriented treatment.

At the core of the book are the mechanisms that make vLLM distinctive: PagedAttention, continuous batching, KV cache design, and scheduler-driven execution. You will learn how request flow, cache allocation, sequence length, prefix reuse, quantized KV storage, and offloading strategies interact to determine concurrency limits and user-visible performance. The book also covers OpenAI-compatible API serving, streaming semantics, realistic benchmarking, and disciplined troubleshooting, so readers can move from conceptual understanding to evidence-based tuning and operational decisions.

The emphasis throughout is on advanced mental models, trade-offs, and production diagnostics rather than introductory walkthroughs. This is a focused guide for readers comfortable with GPU inference, transformer decoding, and performance measurement who want a precise framework for designing, tuning, and operating high-throughput LLM APIs with confidence.

More in Algorithms & Data Structures

Python for Algorithmic Trading : From Idea to Cloud Deployment - Yves Hilpisch
Learning Algorithms : A Programmer's Guide to Writing Better Code - George Heineman
Fundamentals of Data Structures and Algorithms - Elvis C. Foster

RRP $158.00

$122.75

22%
OFF
Python Using GPT-5 and Gemini - Oswald Campesato
Finite Element Mesh Generation, 2e - Daniel S.H.  Lo
Hacker's Delight - Henry Warren

RRP $97.60

$74.75

23%
OFF
Quick Data Structures : Quick Programming - David Matuszek
The Coder Cafe - Teiva Harsanyi

$166.75

Timeless Algorithms : The Seminal Papers - Gary Sutton