Every chapter in this part has been diagnostic. This one spends a budget.
The budget is $50,000 a year, and the first move is the only one that matters: convert it into the unit the vendors price in. Fifty thousand a year is $4,167 a month, and $4,167 a month is $5.71 an hour, every hour, all year. Every decision below falls out of that number.
What $5.71 an hour buys
Rates from earlier chapters, all H100-class, all normalised to one GPU-hour. The right-hand column is what continuous, round-the-clock capacity the whole budget sustains:
| Where | $/GPU-hour | GPU-hours/year | Continuous H100s |
|---|---|---|---|
| Modal, non-preemptible | $11.85 | 4,220 | 0.48 |
AWS p5 on-demand |
$6.88 | 7,267 | 0.83 |
| RunPod serverless | $4.18 | 11,973 | 1.37 |
| Modal, headline (preemptible) | $3.95 | 12,661 | 1.45 |
| RunPod Pod, H100 PCIe | $2.89 | 17,301 | 1.98 |
| AWS 3-year committed | $2.59 | 19,328 | 2.21 |
The same $50,000 buys 4.6× more compute at the bottom of this table than at the top. Not a better GPU — the identical H100, bought differently.
That spread is the entire chapter in one number. At this budget you are not choosing between models or architectures. You are choosing between 0.48 of a machine and 2.21 of them, and the choice is procurement, not engineering.
You cannot buy the instance. You can buy the rate.
The cheapest row looks unreachable, and that intuition is wrong in a useful way.
A whole p5.48xlarge on a three-year plan costs $181,289 a year — 3.63× the entire budget. You cannot buy that machine. But an EC2 Instance Savings Plan is not a machine; it is a commitment to a dollar amount per hour, and the discounted rate applies to whatever usage fits underneath it.
Commit your $5.71 an hour and you get 27.6% of a p5 instance covered continuously — 2.21 H100s at the committed rate. The lot size that locks you out of owning the instance does not lock you out of pricing it.
And here is the part that should change what you do this week. Lock-in is a cost established that AWS refunds a Savings Plan in full only if the hourly commitment is $100 or less, within seven days and the same UTC calendar month.
$5.71 is 5.71% of that ceiling. You have 17.5× of headroom.
The buyer who is structurally inside AWS's only no-questions exit is the small one. A company committing $600 an hour has no undo at all. You have a week to change your mind, at full refund, ten times a year. The conventional advice — small companies shouldn't commit — has it backwards: you are the buyer who can commit safely, precisely because you are small. What you should not do is commit for three years on day one. Commit small, use the return window as a real trial, and scale the commitment as demand proves itself.
The floors that eat the budget before any compute
At $4,167 a month, fixed charges are not rounding errors.
Modal's Team tier is $250/month — $3,000 a year, 6.0% of the budget before a single GPU-second. That is defensible if you need its 50-GPU concurrency; it is 6% of your year if you don't and were on it by default. The Starter tier is $0 with 10 GPU concurrency, which at this budget is not obviously the binding constraint.
RunPod's defaults deserve a second look for a different reason. Serverless GPU versus dedicated sourced them: max workers defaults to 3, and after three days with no requests an endpoint's max workers drops to 2 and stays there until you raise it by hand. Neither costs money. Both quietly cap what your budget can actually deliver on the day traffic arrives.
Storage: where portability is bought
Compute is rented and returned. Data accumulates, and the accumulated data is what makes leaving expensive. This is the one layer where a small buyer should pay deliberate attention, and the numbers are unusually clean.
| Cloudflare R2 | Amazon S3 Standard | |
|---|---|---|
| Storage | $0.015 / GB-month | $0.023 / GB-month |
| Write-class operations | $4.50 / million | $5.00 / million |
| Read-class operations | $0.36 / million | $0.40 / million |
| Egress to the internet | Free | not published in AWS's pricing feed |
Two things fall out.
R2 prices operations at exactly 0.90× S3 — both classes, to the cent. $4.50/$5.00 and $0.36/$0.40 are both precisely 90%. That is a deliberate undercut, not a coincidence of rounding. On storage the gap is wider: R2 is 34.8% cheaper than S3 measured against S3's rate, which is the same fact as S3 being 53.3% more expensive measured against R2's rate. Both directions are stated because the denominator changes the number by half.
But the operations discount is not the point. The egress line is. Cloudflare states plainly that egress "does not incur data transfer (egress) charges and is free" — with a caveat worth reading exactly: that applies to egressing directly from R2 via its Workers API, S3 API and r2.dev domains, and "if you connect other metered services to an R2 bucket, you may be charged by those services."
Free egress is not a discount. It is the purchase of an exit. The cost of moving your data somewhere else later is the single largest component of platform lock-in that a small company actually controls, and one vendor here has priced it at zero.
R2 also exposes an S3 API, which is why this is a portability recommendation and not a vendor recommendation: the same client code addresses both. Cloudflare's Sippy migration tool is "free to use," charging only for the R2-side operations as objects are copied across on demand — though it notes "your source bucket might incur additional charges," which is the egress bill you are leaving behind, payable once.
Two small-scale gotchas that matter more at $50k than at $50m. R2's free tier is 10 GB-month, 1 million write-class and 10 million read-class operations per month — genuinely enough for a small project's metadata and artifacts. And R2 rounds every usage figure up to the next billing unit: 1.1 GB-month bills as 2. At small volumes that rounding is a larger proportion of your bill than it will ever be later.
When you should go hyperscaler anyway
This chapter is not an argument against AWS, and there are cases where the answer is unambiguous.
When your compliance geography demands it. Regions, residency, and sovereignty found H100 capacity in 15 of 106 AWS regions and exactly one region that is both inside the EU and has them. No neocloud in this book publishes a comparable regional map.
When you need audited financial substance from your vendor. The neocloud landscape showed that only CoreWeave and Nebius file audited disclosures; RunPod, Lambda, Crusoe, Modal and Vast.ai are not SEC registrants. If your own customers demand vendor diligence you can evidence, that constraint may decide it for you.
When the rest of your system already lives there. Moving the GPU while leaving the database, the queue and the object store behind buys you a cheaper GPU-hour and a data-transfer bill, and this book cannot tell you the size of the second number.
What I could not price
Internet egress is still unsourced, and I now know precisely why. AWS publishes a machine-readable S3 metered-unit map with 392 keys per region — every storage class, every request type, every retrieval fee. It contains no internet data-transfer key at all; the only outbound entry is intra-region replication at $0.015/GB. So the S3-versus-R2 comparison above is complete on storage and operations and deliberately blank on the line that matters most for leaving. Price your own egress against your own volume before treating any of this as a total.
No AWS, Azure or GCP support-plan floors appear. Support pricing has no metered-unit map (/support returns 404) and the pricing pages render client-side. Those plans are a genuine fixed cost for a small buyer and I could not resolve one.
The diagnostic
- Convert your budget to $/hour before anything else. Annual figures hide the decision; hourly figures make every vendor page directly comparable.
- Is your hourly number under $100? If so, AWS's seven-day full-refund return applies to you and you should use it as a paid trial rather than avoiding commitments entirely.
- Commit small and ratchet, don't commit long. The rate is identical at any commitment size — lock-in is a cost showed twelve published rates with no volume dimension — so there is no discount reward for over-committing, only risk.
- Add up your fixed monthly charges and divide by the budget. One $250/month plan is 6% of a $50k year.
- Check your concurrency defaults today. RunPod ships at 3 max workers and throttles to 2 after three idle days.
- Is your object store's egress free? If not, you are accruing an exit fee every month, and it is the one lock-in cost you can eliminate by choosing differently on day one.
- Are you paying for guarantees you compared against non-guarantees? Modal's headline rate is preemptible; matched guarantees are 3× and move you from 1.45 continuous H100s to 0.48.
What this chapter is not saying
It is not saying be cheap. The 4.6× spread is not a quality gradient — the bottom row of that table is AWS, at its committed rate, and the top row is Modal's guaranteed tier. Cheap and premium are not lined up the way the price suggests.
It is saying that at $50,000 a year the largest lever available to you is procurement, and it is worth more than any optimisation you could make to the model. Nothing in Part 4's serving techniques will find you 4.6×. Choosing where to buy, committing an amount small enough to undo, and putting your data somewhere with no exit toll — those three decisions, made in an afternoon, are the difference between half a GPU and two.