Chapter 3.22 of 8 in this part

The accelerator ladder

Fifteen rungs spanning 13.9× in price. Rank them by dollars per unit of memory bandwidth instead and the order scrambles — the cheapest card on the ladder is also the cheapest bandwidth, and the mid-range card everyone buys is the most expensive thing on it.

8 min read·revised 2026-08-08

Memory bandwidth is the product established that decode is bandwidth-bound. This chapter does the obvious next thing, which almost no vendor page does: takes a real price ladder and divides it by the thing that actually determines throughput.

The result is not a gentle re-ordering. The ladder inverts at both ends.

The ladder, as sold

One vendor, one page, one day — RunPod's Secure Cloud dedicated Pod rates, captured 2026-08-08, alongside the VRAM they list next to each price:

GPU $/hour VRAM
A40 $0.44 48 GB
RTX A6000 $0.53 48 GB
L40 $0.82 48 GB
RTX 6000 Ada $0.84 48 GB
L40S $0.99 48 GB
RTX 5090 $0.99 32 GB
A100 PCIe $1.39 80 GB
A100 SXM $1.49 80 GB
RTX Pro 6000 $1.99 96 GB
H100 PCIe $2.89 80 GB
H100 SXM $2.99 80 GB
H100 NVL $3.19 94 GB
H200 $4.39 141 GB
B200 $5.89 180 GB
B300 $7.39 288 GB

Top to bottom is 13.9× — B300 against the RTX A6000. Against the A40 it is 16.8×.

That span is the fact everyone quotes and the least useful thing on the page. Nobody buys an hour of GPU. They buy the ability to hold a model and stream tokens out of it, and those are the two normalisations that follow.

Divide by bandwidth and the order scrambles

Memory bandwidth from NVIDIA's own specification tables, against the same prices:

GPU $/hour Bandwidth $/hour per TB/s
A40 $0.44 696 GB/s $0.63
A100 PCIe $1.39 1,935 GB/s $0.72
A100 SXM $1.49 2,039 GB/s $0.73
H100 SXM $2.99 3,350 GB/s $0.89
H200 $4.39 4,800 GB/s $0.91
L40S $0.99 864 GB/s $1.15

Read the right-hand column against the left. The cheapest card per hour is the cheapest bandwidth. The most expensive bandwidth on this table costs $0.99 an hour.

The L40S is 81.2% more expensive per TB/s than the A40 — the same fact as the A40 being 44.8% cheaper per TB/s than the L40S, and both directions are stated because the denominator moves the number by half. In absolute terms the L40S is the better chip: newer architecture, far more tensor throughput, AV1 encoders, a 350 W envelope against the A40's 300 W. On the metric that governs decode, it is the worst buy on the ladder.

Notice also how flat the middle is. From A100 PCIe to H200 — a 3.2× jump in list price — bandwidth-per-dollar moves from $0.72 to $0.91. You are paying roughly a quarter more per unit of bandwidth for a chip that costs three times as much, and getting 2.5× more bandwidth in one device. That is not a bad deal; it is a different purchase. More on that below.

Divide by VRAM and it scrambles again

GPU $/hour $/GB-hour
A40 $0.44 $0.0092
RTX A6000 $0.53 $0.0110
L40 $0.82 $0.0171
A100 PCIe $1.39 $0.0174
A100 SXM $1.49 $0.0186
L40S $0.99 $0.0206
RTX Pro 6000 $1.99 $0.0207
B300 $7.39 $0.0257
H200 $4.39 $0.0311
B200 $5.89 $0.0327
H100 NVL $3.19 $0.0339
H100 PCIe $2.89 $0.0361
H100 SXM $2.99 $0.0374

The B300 — the most expensive GPU on the ladder at $7.39 an hour — is cheaper per gigabyte than every H100 variant. It undercuts the H100 SXM by 31.3% measured against the H100's rate, which is the H100 being 45.7% more expensive measured against the B300's. The flagship is not the premium tier on this metric. The H100 tier is.

And the whole-ladder spread on memory is 4.08×: an H100 SXM gigabyte-hour costs 4.08 times an A40 gigabyte-hour.

So we now have three different orderings of the same fifteen products, and the one printed on the pricing page agrees with neither of the other two.

What you are actually paying for

The explanation is not complicated, and memory bandwidth is the product already supplied it: the price of a data-centre GPU tracks its compute, and decode does not use its compute.

The L40S carries 1,466 TFLOPS of tensor throughput and 864 GB/s of memory bandwidth. You are charged for the first number and rate-limited by the second. On the A40 you are charged for much less compute and rate-limited by 696 GB/s — 80% of the L40S's bandwidth for 44% of the price.

This is why the ladder cannot be walked by price. Each rung bundles compute, memory capacity, and memory bandwidth in a different ratio, and only you know which of the three is binding for your workload. Buying "one step up" optimises nothing in particular.

Feasibility first, value second

Everything above is a value calculation, and value calculations are the second question. The first is whether the thing can run at all.

48 GB is 48 GB. The A40 wins both normalised tables and cannot hold a model that needs 80 GB at any price, because you cannot buy 1.7 A40s. Per-unit cheapness only helps when the unit is divisible across devices, and a model that must fit in one GPU's memory is exactly the case where it isn't.

That is what the expensive rungs are really selling. The H200's 141 GB and the B300's 288 GB are not primarily better value — they are the ability to run a configuration the cheaper rungs cannot express, in one device, without the parallelism strategies tax of splitting a model across cards. The B300's excellent $/GB is a consequence of that capacity, not a reason to buy it.

So the sequence is: establish what must fit and what latency you owe, eliminate every rung that cannot meet it, and only then rank the survivors by dollars per unit of the binding resource. Most teams do this in the opposite order and end up on the L40S.

The diagnostic

  1. Which resource binds — capacity, bandwidth, or compute? For decode-dominated serving it is bandwidth. If you have not asked, you are optimising the wrong column.
  2. Does your model fit, in one device, with its KV cache? This is a yes/no gate. Answer it before any price comparison.
  3. Compute $/hour ÷ TB/s for your shortlist. The numbers are all public; the division takes a minute and reorders the list.
  4. Is a newer, better chip actually worse for you? The L40S over the A40 is a real, defensible purchase for anything compute-bound — and an 81.2% overpay per unit of bandwidth if you are decoding.
  5. Have you checked the top of the ladder on $/GB? If capacity is your constraint, the flagship may be the cheap option — the B300 beats every H100 variant per gigabyte.
  6. Are you comparing devices or systems? These are single-GPU rates. Interconnect changes the answer once a job spans cards, and none of it is priced here.

What this chapter is not saying

It is not saying buy A40s. It is saying that a ladder printed in dollars per hour encodes a ranking that is wrong for most inference workloads, and that the correction is one division you can do yourself from public numbers.

The deeper point is that there is no "value" rung. The A40 is the best bandwidth-per-dollar and the best capacity-per-dollar and is unusable for a 70B model at full precision. The B300 is the most expensive thing on the list and the cheapest capacity above 141 GB. Value and feasibility rank the ladder differently, and the pricing page shows you neither.

Sources & methodcaptured 2026-08-08

Sources, captured 2026-08-08. Prices are RunPod's published Secure Cloud dedicated-Pod hourly rates, read from their public pricing page today, along with the VRAM figure RunPod lists beside each GPU — price and capacity therefore come from the same vendor page, which is why the $/GB table uses RunPod's own capacity numbers rather than NVIDIA's. Memory bandwidth and TDP are quoted from NVIDIA's own specification tables: A100 80GB PCIe 1,935 GB/s and 80GB SXM 2,039 GB/s (300 W and 400 W TDP respectively) from the A100 product page; L40S 48GB GDDR6 at 864 GB/s, 350 W and 1,466 TFLOPS tensor performance from the L40S page; A40 48GB GDDR6 at 696 GB/s, 300 W from the A40 page; L4 24 GB at 300 GB/s, 72 W from the L4 page. The H100 SXM 3.35 TB/s and H200 4.8 TB/s figures were captured for memory bandwidth is the product and are reused. Computed by me and verified in a separate pass: the 13.9× and 16.8× price spans, every $/hour-per-TB/s figure, every $/GB-hour figure, the 1.81× bandwidth-value spread and its 81.2%/44.8% two-base statement, the 4.08× VRAM-value spread, the 31.3%/45.7% B300-versus-H100-SXM comparison, and the 3.2× A100-PCIe-to-H200 price step. Both directions are given for every "cheaper/more expensive" claim because the denominator changes these numbers substantially. Bandwidth is not published here for every priced rung. RTX A6000, L40, RTX 6000 Ada, RTX 5090, RTX Pro 6000, H100 PCIe, H100 NVL, B200 and B300 appear in the price and VRAM tables but not in the bandwidth table, because I resolved NVIDIA specification tables for only six of the fifteen and will not mix sourced with recalled figures; the bandwidth table is explicitly a six-rung subset. The claim that the L40S is "the worst buy on the ladder" is scoped to bandwidth-per-dollar and stated as such — on compute-per-dollar it would rank very differently, and no compute-normalised table appears here because I did not resolve comparable tensor-throughput figures across all six. Rates are one vendor on one day; RunPod also publishes a lower Community Cloud tier not used here, and the Secure Cloud tier was chosen for consistency with earlier chapters. These are single-GPU rates and ignore interconnect entirely, which is a real omission for multi-GPU jobs and is deferred to a later chapter rather than estimated. The three-orderings framing, the feasibility-before-value sequence, and the diagnostic are mine.

Want this done on your account rather than by you?

The handbook is the method, written out in full so you can run it yourself — that is the point of publishing it. If you would rather someone else did the first pass, the teardown is free and you keep the findings either way.