Memory bandwidth is the product established that decode is bandwidth-bound. This chapter does the obvious next thing, which almost no vendor page does: takes a real price ladder and divides it by the thing that actually determines throughput.
The result is not a gentle re-ordering. The ladder inverts at both ends.
The ladder, as sold
One vendor, one page, one day — RunPod's Secure Cloud dedicated Pod rates, captured 2026-08-08, alongside the VRAM they list next to each price:
| GPU | $/hour | VRAM |
|---|---|---|
| A40 | $0.44 | 48 GB |
| RTX A6000 | $0.53 | 48 GB |
| L40 | $0.82 | 48 GB |
| RTX 6000 Ada | $0.84 | 48 GB |
| L40S | $0.99 | 48 GB |
| RTX 5090 | $0.99 | 32 GB |
| A100 PCIe | $1.39 | 80 GB |
| A100 SXM | $1.49 | 80 GB |
| RTX Pro 6000 | $1.99 | 96 GB |
| H100 PCIe | $2.89 | 80 GB |
| H100 SXM | $2.99 | 80 GB |
| H100 NVL | $3.19 | 94 GB |
| H200 | $4.39 | 141 GB |
| B200 | $5.89 | 180 GB |
| B300 | $7.39 | 288 GB |
Top to bottom is 13.9× — B300 against the RTX A6000. Against the A40 it is 16.8×.
That span is the fact everyone quotes and the least useful thing on the page. Nobody buys an hour of GPU. They buy the ability to hold a model and stream tokens out of it, and those are the two normalisations that follow.
Divide by bandwidth and the order scrambles
Memory bandwidth from NVIDIA's own specification tables, against the same prices:
| GPU | $/hour | Bandwidth | $/hour per TB/s |
|---|---|---|---|
| A40 | $0.44 | 696 GB/s | $0.63 |
| A100 PCIe | $1.39 | 1,935 GB/s | $0.72 |
| A100 SXM | $1.49 | 2,039 GB/s | $0.73 |
| H100 SXM | $2.99 | 3,350 GB/s | $0.89 |
| H200 | $4.39 | 4,800 GB/s | $0.91 |
| L40S | $0.99 | 864 GB/s | $1.15 |
Read the right-hand column against the left. The cheapest card per hour is the cheapest bandwidth. The most expensive bandwidth on this table costs $0.99 an hour.
The L40S is 81.2% more expensive per TB/s than the A40 — the same fact as the A40 being 44.8% cheaper per TB/s than the L40S, and both directions are stated because the denominator moves the number by half. In absolute terms the L40S is the better chip: newer architecture, far more tensor throughput, AV1 encoders, a 350 W envelope against the A40's 300 W. On the metric that governs decode, it is the worst buy on the ladder.
Notice also how flat the middle is. From A100 PCIe to H200 — a 3.2× jump in list price — bandwidth-per-dollar moves from $0.72 to $0.91. You are paying roughly a quarter more per unit of bandwidth for a chip that costs three times as much, and getting 2.5× more bandwidth in one device. That is not a bad deal; it is a different purchase. More on that below.
Divide by VRAM and it scrambles again
| GPU | $/hour | $/GB-hour |
|---|---|---|
| A40 | $0.44 | $0.0092 |
| RTX A6000 | $0.53 | $0.0110 |
| L40 | $0.82 | $0.0171 |
| A100 PCIe | $1.39 | $0.0174 |
| A100 SXM | $1.49 | $0.0186 |
| L40S | $0.99 | $0.0206 |
| RTX Pro 6000 | $1.99 | $0.0207 |
| B300 | $7.39 | $0.0257 |
| H200 | $4.39 | $0.0311 |
| B200 | $5.89 | $0.0327 |
| H100 NVL | $3.19 | $0.0339 |
| H100 PCIe | $2.89 | $0.0361 |
| H100 SXM | $2.99 | $0.0374 |
The B300 — the most expensive GPU on the ladder at $7.39 an hour — is cheaper per gigabyte than every H100 variant. It undercuts the H100 SXM by 31.3% measured against the H100's rate, which is the H100 being 45.7% more expensive measured against the B300's. The flagship is not the premium tier on this metric. The H100 tier is.
And the whole-ladder spread on memory is 4.08×: an H100 SXM gigabyte-hour costs 4.08 times an A40 gigabyte-hour.
So we now have three different orderings of the same fifteen products, and the one printed on the pricing page agrees with neither of the other two.
What you are actually paying for
The explanation is not complicated, and memory bandwidth is the product already supplied it: the price of a data-centre GPU tracks its compute, and decode does not use its compute.
The L40S carries 1,466 TFLOPS of tensor throughput and 864 GB/s of memory bandwidth. You are charged for the first number and rate-limited by the second. On the A40 you are charged for much less compute and rate-limited by 696 GB/s — 80% of the L40S's bandwidth for 44% of the price.
This is why the ladder cannot be walked by price. Each rung bundles compute, memory capacity, and memory bandwidth in a different ratio, and only you know which of the three is binding for your workload. Buying "one step up" optimises nothing in particular.
Feasibility first, value second
Everything above is a value calculation, and value calculations are the second question. The first is whether the thing can run at all.
48 GB is 48 GB. The A40 wins both normalised tables and cannot hold a model that needs 80 GB at any price, because you cannot buy 1.7 A40s. Per-unit cheapness only helps when the unit is divisible across devices, and a model that must fit in one GPU's memory is exactly the case where it isn't.
That is what the expensive rungs are really selling. The H200's 141 GB and the B300's 288 GB are not primarily better value — they are the ability to run a configuration the cheaper rungs cannot express, in one device, without the parallelism strategies tax of splitting a model across cards. The B300's excellent $/GB is a consequence of that capacity, not a reason to buy it.
So the sequence is: establish what must fit and what latency you owe, eliminate every rung that cannot meet it, and only then rank the survivors by dollars per unit of the binding resource. Most teams do this in the opposite order and end up on the L40S.
The diagnostic
- Which resource binds — capacity, bandwidth, or compute? For decode-dominated serving it is bandwidth. If you have not asked, you are optimising the wrong column.
- Does your model fit, in one device, with its KV cache? This is a yes/no gate. Answer it before any price comparison.
- Compute $/hour ÷ TB/s for your shortlist. The numbers are all public; the division takes a minute and reorders the list.
- Is a newer, better chip actually worse for you? The L40S over the A40 is a real, defensible purchase for anything compute-bound — and an 81.2% overpay per unit of bandwidth if you are decoding.
- Have you checked the top of the ladder on $/GB? If capacity is your constraint, the flagship may be the cheap option — the B300 beats every H100 variant per gigabyte.
- Are you comparing devices or systems? These are single-GPU rates. Interconnect changes the answer once a job spans cards, and none of it is priced here.
What this chapter is not saying
It is not saying buy A40s. It is saying that a ladder printed in dollars per hour encodes a ranking that is wrong for most inference workloads, and that the correction is one division you can do yourself from public numbers.
The deeper point is that there is no "value" rung. The A40 is the best bandwidth-per-dollar and the best capacity-per-dollar and is unusable for a 70B model at full precision. The B300 is the most expensive thing on the list and the cheapest capacity above 141 GB. Value and feasibility rank the ladder differently, and the pricing page shows you neither.