ANAlpesh Nakrani
SolutionsBlogBooksPraiseAboutWork with me
All books
Multimodal in Practice cover
Cover preview Page 1 of 150
Free PDF/Technical Deep Dives

Multimodal in Practice

Building AI Systems That See, Hear, Read, and Still Fail in New Ways

A book by Alpesh Nakrani

Subscribe to get the complete book. Your PDF downloads immediately.

Pages
150
Chapters
16
Format
PDF · 27.6 MB
01/Overview

What's in the book

Adding images and audio to a system is not a parameter change. It is a new set of failure modes. This deep dive covers what multimodal models actually perceive, where they hallucinate differently, and the evaluation that catches it.

Vision and audio change the failure modes, not just the inputs. What breaks when the model has to see and hear.

MultimodalEngineering
02/Inside the PDF

Table of contents

PDF page
  1. FMFront Matter: Multimodal in Practice
  2. INTIntroduction: The Photo of the Bumper17
  3. 01The Demo Is Not the System23
  4. 02Every Modality Brings Its Own Blind Spots31
  5. 03A Vision Answer Is Not Evidence39
  6. 04Encoders, Contrastive Learning, and Cross-Modal Alignment47
  7. 05Joint Embedding Spaces and Their Limits54
  8. 06Documents Are Multimodal Objects61
  9. 07OCR Is Not Document Understanding68
  10. 08Grounding, Bounding Boxes, and Visual Citations75
  11. 09Audio Is Not Text After Transcription82
  12. 10Video Is Not Just Frames89
  13. 11Multimodal RAG and Cross-Modal Search96
  14. 12Evaluation: Seeing Is Not Verifying103
  15. 13Production Architecture for Multimodal Systems110
  16. 14Cost, Latency, and the Price of Vision Tokens117
  17. 15Privacy, Safety, and High-Stakes Images124
  18. 16The Use Case Field Guide131
  19. AAppendix A: Back Matter138

The full chapters, illustrations, and reference material are included in the downloadable PDF.

Read it at your own pace

Your next read, ready to download.

150 pages. The complete edition. Free with your subscription.

Download free PDF

Ask AI about Multimodal in Practice — Free PDF