
The Economics of Inference
Shipping models you can afford to run
A book by Alpesh Nakrani
Subscribe to get the complete book. Your PDF downloads immediately.
- Pages
- 92
- Chapters
- 11
- Format
- PDF · 17.6 MB
What's in the book
The model that wins is the one you can afford to run at scale and explain to a customer. This manual works through quantization, routing, caching, and the plain arithmetic that decides whether an AI feature has a margin or just a demo.
Revenue rarely rewards the biggest model. A working account of latency, cost, and the smaller model that wins the P&L.
Table of contents
PDF page- INTIntroduction: The Request That Won on Quality and Lost on Margin9
- 01The Arithmetic Nobody Runs Before Launch15
- 02Gross Margin Per Completed Task Is the Only Metric That Matters21
- 03Routing to the Cheapest Model That Still Clears the Bar27
- 04Caching the Request You Already Answered33
- 05Batching and the Shape of Your Traffic39
- 06Quantization and the Margin You Compress Into45
- 07Fallback and Cascade Design51
- 08Cost Attribution and Observability58
- 09The Build-Versus-Buy Cost Curve65
- 10Pricing the Feature to the Customer71
- 11Where Premium Capability Earns Its Keep77
- ENDConclusion: Price Every Call Before It Becomes a Habit83
The full chapters, illustrations, and reference material are included in the downloadable PDF.

The Economics of Inference
The complete edition,
ready to read offline.
Get your free copy.
Subscribe to Alpesh’s book and writing updates. Unsubscribe anytime.
Your download is ready.
You’re subscribed. Your PDF download should start automatically.
Download PDF againIf it doesn’t start, use the button above.



