Get Free Shipping on orders over $79
Practical LLM Evaluation for Production Systems : Measure, monitor, and improve AI system reliability across training and inference - Ammar Mohanna
eTextbook alternate format product

Instant online reading.
Don't wait for delivery!

Go digital and save!

Practical LLM Evaluation for Production Systems

Measure, monitor, and improve AI system reliability across training and inference

By: Ammar Mohanna, Indrajit Kar, Zonunfeli Ralte

Paperback | 30 June 2026

At a Glance

Paperback


$123.75

or 4 interest-free payments of $30.94 with

 or 

Ships in 10 to 15 business days

Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applications

Free with your book: DRM-free PDF version + access to Packt's next-gen Reader*

Key Features:

- Design evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systems

- Measure quality, safety, grounding, robustness, and production readiness with practical metrics

- Apply unified evaluation methods to text, multimodal, and agentic AI systems

Book Description:

Modern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.

Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.

Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.

By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.

*Email sign-up and proof of purchase required

What You Will Learn:

- Design repeatable evaluation pipelines for LLM systems

- Assess inference quality, latency, and operational cost

- Evaluate multimodal, agentic, and reasoning AI systems

- Build regression gates and deployment evaluation workflows

- Detect hallucinations and grounding failures in VLMs

- Assess routing stability in mixture-of-experts models

- Evaluate Text2SQL, OCR, and retrieval-based systems

- Translate evaluation signals into production decisions

Who this book is for:

ML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts.

Table of Contents

- Foundations of LLM Evaluation: Core Concepts and Primitives

- Building Reliable Text-Only LLMs Through Training-Time Evaluation

- Controlling Text-Only LLM Behavior at Inference Time

- Grounding and Reliability in Vision Language Models During Training

- Evaluating Visual Grounding and Reliability at Inference Time

- Evaluating Multimodal Conversational LLMs Across Training and Inference

- Evaluating Routing and Reliability in Mixture of Experts LLMs

- Evaluating Reliability and Control in Computer-Using Agent Systems

- Evaluating Information Extraction and Document-Understanding LLMs

- Evaluating Reasoning LLMs in Depth

- Evaluating Specialized LLM Systems

More in Databases

Statistics and Data Handling for Biologists : A Student's Guide - Neil Millar
Dynamics of Marine Structures - Yingguang  Wang

RRP $503.95

$442.99

12%
OFF
Social Research Methods : 4th Edition - Maggie Walter

RRP $101.95

$87.75

14%
OFF
Data Empire : How information shaped human history - Roopika Risam
Python All-in-One For Dummies : 3rd Edition - John C. Shovic

RRP $74.95

$55.75

26%
OFF
Building a Scalable Data Warehouse with Data Vault 2.0 - Dan Linstedt
Fundamentals of Database Systems, Global Edition : 7th edition - Ramez Elmasri
DAMA-DMBOK : Data Management Body of Knowledge - DAMA International

RRP $137.49

$106.75

22%
OFF
Data Analytics for Accounting ISE : 3rd Edition - Vernon J. Richardson

RRP $169.95

$146.75

14%
OFF