Library
Reading / 1 min read

The bookshelf — five books worth owning for AI and cloud cost

Published August 4, 2026
Summary

The five physical books worth owning in a field that doesn't have its own book yet — one per layer of the stack, from FinOps canon to serving systems — plus the sixth you assemble yourself from the primary literature.

The five physical books worth owning in a field that doesn't have its own book yet — one per layer of the stack.

1. Cloud FinOps, 2nd ed. — Storment & Fuller (O'Reilly, 2023)

The canon of the discipline this field extends. Read it to learn it — and to study the seams, because it barely mentions tokens.

2. AI Engineering — Chip Huyen (O'Reilly, 2025)

The best modern application-layer textbook: model selection, evals, inference optimization. An engineer's book, in an engineer's language.

3. Chip War — Chris Miller (Scribner, 2022)

The silicon layer as narrative: TSMC, fab economics, why compute is geopolitically scarce. The history under every GPU price.

4. Large Language Model-Based Solutions — Shreyas Subramanian (Wiley, 2024)

The only book in print explicitly about LLM cost-effectiveness. It predates the agent era — its gaps are the field's open questions.

5. Hands-On LLM Serving and Optimization — Wang & Hu (O'Reilly, 2026)

Serving systems in book form: batching, quantization, deployment economics. Pair with the weekend-H100 lab.

The sixth book you assemble yourself

Print the canonical papers — Kipply, Chinchilla, the Epoch economics papers, Mooncake, DistServe, FrugalGPT, the FinOps AI working-group papers — into a binder and annotate by hand. In a field this young, the primary literature is the textbook.

For a scored, layered reading list of the wider shelf, see the reading stack in the Engineer's Handbook.

Keep reading

EssayAugust 4, 20263 min read

Why I don't do pay-from-savings — and what I do instead

Contingency pricing rewards the easy 10% of cloud savings and structurally ignores the hard 30%. Why a fixed fee, agreed before the work starts, is the only model that keeps an advisor independent — and what the engagement includes instead.

Cost HandbookAugust 11, 20266 min read

Attribution at scale

Attribution fails at the join, not the tag. The three clouds disagree about case, length and count, so the only portable key is the intersection — 63 characters, lowercase — and on the trace side the GenAI token convention is still marked Development.

Cost HandbookAugust 11, 20266 min read

Capacity planning and forecasting

Forecasting dollars for AI is forecasting a moving target. Forecast capacity instead — and note that a commitment discount is really a utilisation threshold. At AWS's published 62.4% three-year discount, the committed rate is 37.6% of on-demand, so the machine has to be busy 37.6% of the time to break even.