AI hardware ROI for a used-GPU inference build: a sensitivity analysis that can change the answer
A second-hand accelerator comes with no warranty, no service history, and no way to know how hard the last owner ran it. That's the risk.
What the used price tag doesn't cover
The trade: older cards, a case, board, memory, PSU, and cooling that can feed them, plus a plan for when one dies. The fit: steady local inference where memory capacity beats this year's architecture, you're buying VRAM per dollar. Weigh it against new parts, a refurbished workstation, or paying per token through an API, and price one unit throughout: a quality-gated inference hour over whatever's left of the card's life, not tokens or GPU-hours. The point is showing which guess controls the decision, not a tidy verdict, so use live prices and rates: they age faster than the logic.
The measurements the model actually eats
Split cost into fixed and variable. Card, rig, and reserved capacity are fixed; electricity, paid tokens, and data transfer move with the work. A dedicated operator is fixed; per-exception patching is variable, since it scales with failures. Subtract anything you'd buy anyway: a laptop you own isn't an AI cost, a memory upgrade bought only for the model is. Run the same workload down local, rented, and API routes and log:
- accepted tasks against severe failures
- retries, abstentions, minutes of human correction
- queue, first-token, and p95 latency
- active, idle, sleeping, and dead hours
- energy draw, cooling load, maintenance time
Quality belongs in the price: two tries and a cleanup pass get priced on the finished result, not the first attempt, and a local specialist nailing a narrow job shouldn't be benchmarked against a frontier API at full effort. Used hardware carries its own risk too: surprise platform upgrades, weak power draw, resale that's hard to unload. Model demand as a range: a maxed-out box has no room for an interactive request, and batch work counts only if you need it. The recurring error: one precise payback month from soft inputs. Date the volatile ones; don't hide uncertainty in decimal places.
Payback is a boundary, not a date
Payback month is the first month cumulative discounted benefit passes cost; ROI is net discounted benefit over discounted cost. Both need a conservative useful life and a residual value, since hardware can outlive the model family or its software support. Productivity gains need a haircut: ten minutes saved per developer isn't ten minutes of revenue. Count it only if it avoids a hire, cuts outsourced spend, or shortens a binding queue, then rerun the model with that value at zero.
Weigh hourly rental against buying outright: a short benchmark is cheap insurance against a bad purchase, and it prices the storage, setup, and idle instances the headline rate hides, the same blind spot behind a monthly plan priced wrong for pay-per-use work. State the boundary in operational terms: tasks per month, GPU hours, max API cost per task. Set a date to recheck when volume, pricing, or support status moves. ROI is a relationship between one workload, one alternative, and dated assumptions, not a fixed number.
None of that fixes the real failure mode: buying the card first and building the spreadsheet to agree with you after. A worksheet catches a bad demand curve fine. It has no defense against wanting the thing before you opened it.