AI hardware ROI for a used-GPU inference build: comparing the complete purchase price
Buy the used card for the VRAM, not the model number, and assume the checkout price is missing most of what you'll actually pay. That's the whole claim. The qualifier: a used-GPU build still makes sense, the sticker price is just a rounding error in the number that matters.
This is second-hand accelerators in a matching case, board, memory, PSU, and cooling, running a workload that stays busy for months, not spiking once. Memory capacity beats newer silicon here, so dollars per gigabyte of VRAM can be excellent if the card holds up. Compare it against new consumer hardware, a refurbished workstation, or paying by the token, the same full cost architecture question in different clothes. What you're pricing isn't tokens or GPU-hours, it's a quality-gated inference hour over whatever life the hardware has left.
the real bill of materials
Installed cost is compute plus host plus memory plus storage plus power plus cooling plus network plus tax plus setup. Add it up, build the same bill for the alternative, then strip what you'd have bought anyway: your daily laptop doesn't count, a memory upgrade bought solely for a bigger model does. Split what's left into fixed and variable. Hardware, installation, and reserved capacity sit still once paid for; electricity, tokens, transfer, and per-incident review move with usage, and a dedicated operator is a fixed salary while someone patching exceptions is variable, scaling with your failure rate.
Then measure the workload you actually run, not the demo: accepted tasks against severe failures, input, cached input, reasoning, and output tokens, how often a result needs a retry or a human hand, latency out to the p95 case, and hours active, idle, or asleep. A model needing two attempts and five minutes of cleanup gets priced on the completed result, not the fast first draft. Skip that and you're comparing a cheap wrong answer to an expensive right one.
where the estimate actually breaks
Second-hand hardware carries unknown wear, no warranty, hidden platform upgrades, higher power draw, downtime, and poor resale value. Model low, base, and high demand instead of one number, and respect the p95 case: a fully utilized box has nothing left for an interactive request that lands wrong. Idle states matter too: a workstation asleep overnight differs from a server holding models resident for instant access, and on shared systems, queueing, abandoned requests, and edge-fleet maintenance routinely outweigh electricity. The mistake I see most is pricing a bare card against a finished machine or a turnkey API, which flatters whichever side got the incomplete treatment.
Payback is the first month cumulative discounted benefit clears cumulative cost; ROI is net discounted benefit over discounted cost, both against a conservative useful life and residual value, since hardware can outlive the model it serves. Productivity gains need a realization factor: run the model once with that value at zero, since time savings alone rarely justify a purchase. Price rental too, the same question behind local-first cascade: it earns its keep under uncertain demand, but only if you count storage, transfer, minimum billing, and the instance somebody forgot to shut down. State the break-even in operational terms, accepted tasks per month or minimum useful life, and set a retest date tied to volume, model quality, API price, or electricity cost.
The one rule I'd keep if I dropped everything else: price the incremental cost of the whole system against the workload it actually replaces, never the part against the whole.