
Generative AI Engineering for Developers
Designing Scalable Applications with Large Language Models
By: Steve Millan
eBook | 15 September 2026
At a Glance
ePUB
eBook
$5.54
or 4 interest-free payments of $1.39 with
Instant Digital Delivery to your Kobo Reader App
Large language models are no longer a research curiosity — they are production infrastructure. But the gap between a compelling demo and a system that serves millions of users reliably, economically, and safely is vast. This book closes that gap.
Generative AI Engineering for Developers is a comprehensive, code-first guide for software engineers who want to build production-grade applications powered by LLMs. It goes beyond "how to call the API" to cover the full engineering stack: architecture patterns, evaluation discipline, cost management, observability, security, and the operational practices that separate demos from real products.
Across fifteen chapters and four parts, you will learn how transformers actually work — not at the level of academic mathematics, but at the level of engineering intuition needed to make good architectural decisions. You will master the core application patterns that recur across nearly every LLM product: prompt engineering, retrieval-augmented generation, tool use, and autonomous agents.
You will build the production infrastructure that makes these systems trustworthy: evaluation pipelines, observability stacks, cost controls, and security defenses. And you will explore the advanced techniques — fine-tuning, multimodal systems, and the evolving frontier — that unlock the next level of capability.
What you will learn:
- How the transformer architecture shapes the latency, cost, and quality properties of the systems you build
- Prompt engineering as a rigorous software discipline — not guesswork, but systematic design and measurement
- Retrieval-Augmented Generation: hybrid retrieval, reranking, and building knowledge-grounded applications
- Tool use and function calling patterns for extending models beyond their training knowledge
- Agent architectures — ReAct, multi-agent orchestration, memory systems, and reliability controls
- The economics of inference: model routing, caching strategies, and cost optimization at scale
- Evaluation frameworks that catch regressions before users do
- Observability and monitoring for systems where silent quality failures are the primary threat
- Prompt injection, data exfiltration, and a defense-in-depth security strategy
- Fine-tuning with LoRA, DPO, and parameter-efficient methods — and when not to fine-tune
- Multimodal engineering: document understanding, voice pipelines, and cross-modal retrieval
Who this book is for:
This book is written for software engineers who already understand distributed systems and want to build seriously with LLMs — not just prototype. If you are comfortable with Python and cloud infrastructure and are ready to think about LLM applications the way you think about any production system — with rigor around reliability, cost, observability, and testing — this book was written for you.
About the approach:
Every chapter is grounded in working code. Architectural recommendations are backed by engineering reasoning, not vendor marketing. The patterns described here have been validated in production systems. And because the field moves faster than any book can, the emphasis throughout is on the underlying principles that remain stable even as models, APIs, and benchmarks evolve.
"The engineers who approach LLM applications with curiosity, systematic evaluation, operational discipline, and clear thinking about what these systems can and cannot do will build things that matter."
on
ISBN: 6610001374167
Published: 15th September 2026
Format: ePUB
Language: English
Publisher: PublishDrive
























