Every other chapter in this part compares one rental price to another. This one compares renting to owning, which requires four numbers most people guess: how long the hardware lasts, how much power it draws, what the facility costs, and what the humans cost.
Three of those are sourceable from primary documents. The fourth is not, and pretending otherwise is how self-hosting business cases get written.
Start with depreciation, because it is disclosed
Public companies must tell the SEC how long they assume their servers last. The three largest buyers of this hardware disagree — and one of them changed its mind, in writing, because of AI.
| Filer | Disclosed useful life |
|---|---|
| Alphabet (FY2025 10-K) | "We depreciate servers and network equipment generally over a period of six years." |
| Microsoft (FY2026 10-K) | "servers and network equipment, two to six years" |
| Amazon (FY2025 10-K) | six years → five years, effective 1 January 2025 |
Amazon's language is the one to sit with:
Effective January 1, 2025 we changed our estimate of the useful lives of a subset of our servers and networking equipment from six years to five years. The shorter useful lives are due to the increased pace of technology development, particularly in the area of artificial intelligence and machine learning.
And the cost of that revision, in the same filing: an increase in depreciation and amortization expense of $1.4 billion and a reduction in net income of $1.0 billion, "which primarily impacted our AWS segment."
Read that as an operator, not an investor. The largest buyer of AI hardware on earth looked at its own fleet, concluded the useful life was shorter than it had assumed, and took a billion-dollar hit to say so. If you are amortising a GPU purchase over five years, you are matching Amazon's revised estimate — after the revision. If you are using six, you are using the number Amazon just abandoned. And Microsoft's disclosed range starts at two years, which is a very different business case from six.
Nobody outside your company will tell you which end of that range your hardware sits at. But the range itself is public, audited, and free.
The nameplate is not the draw
Now power, where there is a trap this book has already named in a different context.
NVIDIA publishes two numbers for a DGX H100. The power supply configuration is six 3.3 kW units — 19.8 kW of nameplate, arranged "for 4+2 redundancy," so only 13.2 kW is load-bearing. And in the environmental specifications, the number nobody quotes:
Heat Output: 38,557 BTU/hr
A computer converts essentially all the electricity it draws into heat. So that heat figure is the electrical draw, in different units: 38,557 BTU/hr ÷ 3,412.14 = 11.30 kW.
| Figure | Value | vs actual |
|---|---|---|
| PSU nameplate (6 × 3.3 kW) | 19.8 kW | 1.75× |
| Load-bearing PSUs (4 × 3.3 kW) | 13.2 kW | 1.17× |
| Heat-derived actual draw | 11.30 kW | 1.00× |
Size your colocation contract off the nameplate and you buy 75% more power than the machine uses. At $150–250 per kW per month, that is real money spent on capacity that never draws current.
This is the same rule benchmarking tokens per second per dollar found in MLPerf's rulebook, which explicitly refuses "a TDP configuration, a power supply rating" as a substitute for power "measured at the wall." MLPerf wrote that rule to stop vendors overstating efficiency. The identical error, pointed the other way, makes buyers overprovision facilities.
One more decomposition worth having. Eight H100 SXM at "Up to 700W (configurable)" is 5.6 kW of GPU. Against 11.30 kW of system:
The GPUs are 49.6% of the box's power. The other half is CPUs, 2 TB of DRAM, eight NICs, storage, fans, and conversion losses. Any model that prices "GPU watts" and stops has understated the electricity bill by roughly a factor of two.
What the machine has to cost
Here is the move that keeps this chapter honest: NVIDIA does not publish a DGX list price, so I am not going to invent one. Instead, invert the question. Given what renting costs and what running costs, what would the purchase price have to be for owning to win?
The rental baseline for the same eight H100s, resolved from AWS's pricing feeds today: p5.48xlarge is $55.04/hour on demand and $20.69504/hour on a three-year all-upfront EC2 Instance Savings Plan. Over three years of continuous availability:
- On-demand: $1,446,451
- Three-year committed: $543,866
Owning must beat $543,866 — the committed number, because that is the price a buyer with three years of conviction can actually get. Subtract the running costs and what remains is the capex ceiling:
| Scenario | Power | Colo | Engineering | Capex ceiling |
|---|---|---|---|---|
| $0.08/kWh, PUE 1.2, $150/kW-mo | $28,508 | $61,020 | $0 | $454,338 |
| …with one engineer at $50k/yr | $28,508 | $61,020 | $150,000 | $304,338 |
| $0.12/kWh, PUE 1.4, $250/kW-mo | $49,890 | $101,699 | $0 | $392,276 |
| …with one engineer at $50k/yr | $49,890 | $101,699 | $150,000 | $242,276 |
Look at what the engineering row does. Adding a single part-time engineer at $50k/year consumes $150,000 of the three-year budget — a third of the entire capex ceiling in the favourable scenario.
That is the whole argument. Not power, which is modest. Not colocation, which is predictable. The largest single swing factor in the buy-versus-rent decision is the cost of the person who keeps the machine running, and it is the one line every self-hosting spreadsheet leaves at zero.
The term nobody can source
I could not find a citable figure for the engineering cost of operating GPU infrastructure, and I do not believe one exists in a usable form. Salaries vary by market; the fraction of a person a cluster consumes varies by scale, by how much of the stack you run, and by how good your day is.
So rather than manufacture a number, name the work, because the list is what people underestimate:
- Driver and firmware currency — CUDA, kernel, NIC firmware, and the version matrix between them
- Hardware failure — GPUs fail; RMA cycles are weeks; a dead node is capacity you paid for and cannot use
- Thermal and power incidents — a DGX H100 wants 5–30 °C and moves 1,105 CFM; facility problems become your problems
- Scheduling and multi-tenancy — everything multi-tenancy and co-location describes, but you operate it
- On-call — someone's phone rings at 3am, and that is a real cost even in months when it doesn't
- Opportunity cost — the highest-leverage engineer available is usually the one who would run this
Whatever number you put in that row, put a number in it. A zero there is not a conservative assumption; it is the assumption that produces the wrong answer.
What ownership actually buys
The arithmetic above is not a verdict against owning. Three things it cannot capture, all of which favour buying:
Utilisation above the rental break-even. Commitment maths established that break-even utilisation is the complement of the discount. Owned hardware has no discount ceiling — at genuinely continuous load, the marginal hour costs only electricity.
Capacity that exists. Regions, residency, and sovereignty found p5 in 15 of 106 AWS regions and one EU region. If your compliance geography has no rentable H100, ownership is not the cheap option; it is the only option.
No capacity risk. A Savings Plan is a price commitment, not a machine. Owned hardware is a machine.
Against that, ownership concentrates the risk that Amazon's filing describes: you hold the depreciation exposure on hardware whose useful life the market is revising downward. Renting transfers that risk to someone whose auditors make them disclose it.
The diagnostic
- What useful life are you assuming, and can you defend it against Amazon's five years? This single input moves the annual cost more than anything else on the page.
- Did you size power from the nameplate or the heat spec? On a DGX H100 that is a 1.75× difference in what you buy.
- Did you count non-GPU power? The GPUs are about half the draw.
- What number is in the engineering row? If it is zero, the model is wrong, and by more than power and colo combined.
- Compare against the committed rental rate, not on-demand. $543,866 versus $1,446,451 over three years — using the wrong baseline makes owning look 2.7× better than it is.
- Is your utilisation genuinely continuous? Owned hardware bills you for every hour whether you use it or not, exactly like a commitment, with none of the exit options.
- Does the hardware you need exist to rent in your region? If not, this comparison is academic.
What this chapter is not saying
It is not saying don't buy. At sustained high utilisation, at scale, with staff you already employ and a facility you already have, owning wins clearly — which is why every hyperscaler does it.
It is saying that the self-hosting case is usually built by comparing a purchase price to an on-demand rental rate, and both halves of that comparison are wrong. The rental side should be the committed rate. The ownership side should include a depreciation schedule you can defend, power measured rather than nameplated, and a non-zero number for the humans. Do those three things and the answer may still be buy — but it will be an answer rather than a hope.