Get Free Shipping on orders over $89
Vision Language Models : Building Vlms with Hugging Face - Andres Marafioti
eTextbook alternate format product

Instant online reading.
Don't wait for delivery!

Go digital and save!

Vision Language Models

Building Vlms with Hugging Face

By: Andres Marafioti, Orr Zohar, Miquel Farre, Merve Noyan

Paperback | 14 July 2026

At a Glance

Paperback


$134.75

or 4 interest-free payments of $33.69 with

 or 

Available: 14th July 2026

Preorder. Will ship when available.

Vision language models (VLMs) combine computer vision and natural language processing to create powerful systems that can interpret, generate, and respond in multimodal contexts. Vision Language Models is a hands-on guide to building real-world VLMs using the most up-to-date stack of machine learning tools from Hugging Face, Meta (PyTorch), NVIDIA (Cuda), and others, written by leading researchers and practitioners Merve Noyan, Miquel Farr©, Andr©s Marafioti, and Orr Zohar. From image captioning and document understanding to advanced zero-shot inference and retrieval-augmented generation (RAG), this book covers the full VLM application and development lifecycle. Designed for ML engineers, data scientists, and developers, this guide distills cutting-edge VLM research into practical techniques. Readers will learn how to prepare datasets, select the right architectures, fine-tune and deploy models, and apply them to real-world tasks across a range of industries. Explore core model architectures and alignment techniquesTrain and fine-tune VLMs with Hugging Face, PyTorch, and othersDeploy models for applications like image search and captioningImplement advanced inference strategies, from zero-shot to agentic systemsBuild scalable VLM systems ready for production use

More in Natural Language & Machine Translation

How To Think About AI : A Guide For The Perplexed - Richard  Susskind

RRP $25.95

$22.75

12%
OFF
AI Engineering : Building Applications with Foundation Models - Chip Huyen
AI ChatBots For Dummies : For Dummies (Computer/Tech) - Kelly Noble Mirabella
The AI Engineering Bootcamp : Build, Ship, Share - Greg Loughnane

RRP $107.95

$75.75

30%
OFF
Acting : Keywords and Concepts - John  Matthews

RRP $39.99

$38.75

Acting : Keywords and Concepts - John  Matthews

RRP $130.00

$118.75

Google Gemini For Dummies - Bonaventura Di Bello
AI for Marketing : The Consumer Perspective - Idil M. Cakim

RRP $273.00

$236.99

13%
OFF
Prompt Cartography : Interactive Web Map Design with LLMs - Ian  Muehlenhaus