
2025/Free online book · Field Manuals
The Economics of Inference
Shipping models you can afford to run
Access
Free
Chapters
11
Read time
111 min
The model that wins is the one you can afford to run at scale and explain to a customer. This manual works through quantization, routing, caching, and the plain arithmetic that decides whether an AI feature has a margin or just a demo.
Revenue rarely rewards the biggest model. A working account of latency, cost, and the smaller model that wins the P&L.
This edition is free to read onsite. Each chapter has its own URL, so readers can bookmark, share, and return to the exact section they need.
Table of contents
INTIntroduction: The Request That Won on Quality and Lost on MarginA feature can be impressive and still lose money on every successful request, and nobody notices until finance asks why growth costs more each month.9 min01The Arithmetic Nobody Runs Before LaunchToken pricing looks like a footnote until you multiply it by a million requests, and by then the unit economics are already the product.9 min02Gross Margin Per Completed Task Is the Only Metric That MattersAccuracy, latency, and delight all sound like product metrics until you realize none of them tell you whether the feature makes money.8 min03Routing to the Cheapest Model That Still Clears the BarMost requests are easy, and sending them to the flagship model is a tax you chose to pay on purpose.9 min04Caching the Request You Already AnsweredThe cheapest inference call is the one you never make, and most systems throw that discount away by accident.8 min05Batching and the Shape of Your TrafficLatency-sensitive and latency-tolerant work are different businesses wearing the same API, and batching only pays off once you tell them apart.9 min06Quantization and the Margin You Compress IntoShrinking a model's footprint shrinks the bill, and the honest question is what quality you quietly shrank along with it.8 min07Fallback and Cascade DesignA cascade is a bet that a cheap model's doubt is cheaper to check than a flagship model's certainty is to buy outright.9 min08Cost Attribution and ObservabilityIf you cannot say which customer, feature, or workflow spent the money, you do not have a cost problem, you have a visibility problem wearing one.9 min09The Build-Versus-Buy Cost CurveSelf-hosting looks cheaper on a spreadsheet and looks different once you price the GPUs, the on-call rotation, and the utilization you actually get.8 min10Pricing the Feature to the CustomerThe price on your pricing page and the cost on your inference bill are two different documents, and the gap between them is your margin, on purpose or by accident.8 min11Where Premium Capability Earns Its KeepSome requests deserve the most expensive model you have, and the discipline is knowing which ones without guessing.9 minENDConclusion: Price Every Call Before It Becomes a HabitThe margin does not fail all at once; it leaks one unpriced habit at a time, and the fix is pricing every call before the habit sets in.8 min
