AI hardware ROI for a used-GPU inference build: valuing productivity without inventing savings
I've watched people build the same rig twice without noticing: secondhand accelerators, a board and case that took two tries to fit, no memory budget because "the cards have enough VRAM." Months later, when someone asks whether it paid for itself, nobody has an answer, good or bad: no baseline timing, no logged tasks, power folded into "office costs" months back. That's the failure this piece means to keep you out of: buying capacity on a promise of saved time, never pricing what it was worth.
The workload assumption matters more than the parts list: steady local inference where memory capacity buys more than next-gen silicon, the same local-first cascade call other self-hosting decisions come down to. Alternatives: new hardware, refurbished workstation gear, per-token API calls; used cards promise lower entry price and strong VRAM per currency unit, if the silicon holds. Fix the unit you're buying: a quality-gated inference hour delivered over what's left of the card's life. Tokens and GPU-hours are gauge readings, not the destination.
The rule underneath it: a productivity claim needs measured, verified completion time and an explicit economic use for the freed time. Put your own quotes, electricity rate, and API prices in a worksheet, not fixed numbers in prose. Architecture ages slowly; prices move by the week.
price the hour, not the token
Start from one relationship; the rest is commentary on it:
monthly_value = accepted_tasks * minutes_saved * loaded_hourly_value / 60 * realization_rate
Time comparable work with and without the system, include review, apply a realization rate: saved minutes mostly become slack, not billed output. Split costs. Fixed: hardware, install, reserved capacity. Variable: electricity, paid tokens, transfer, some support. Review sits on the fence: a dedicated operator is fixed, exception-only review is variable, blurring the two flatters your preferred number.
List every incremental line: host, memory, storage, network, power, cooling, enclosure, spares, tax, shipping, install. Subtract what the organization would buy anyway. A laptop someone already owns isn't an AI cost; a memory stick bought purely so a model fits is.
Run candidate local, rented, and API routes side by side and record:
- accepted tasks, severe failures
- input, cached input, reasoning, output tokens
- retries, abstentions, repair minutes
- queue, time to first token, completion, p95
- active, resident-idle, sleep, unavailable hours
- wall energy, cooling share
- maintenance, incident time
A quality adjustment matters: a cheap route needing two attempts and five minutes of correction carries the cost of the completed result, not the first draft. A small local model that reliably clears one narrow job shouldn't be benchmarked against a frontier API at max reasoning effort it never needed.
the used-card discount is a risk premium, not free money
A secondhand rig loses what a warranty would normally cover: unseen wear, drivers that stop updating, an unbudgeted platform upgrade, rising power draw, downtime hunting a replacement, a thin resale market. Model demand three ways, low, base, high, and don't skip p95 latency: a machine at full utilization has nothing left for an interactive request. Batch work can fill quiet hours, only if the organization needs it done.
Meter idle states separately: a box sleeping overnight costs differently than a server holding several models resident against a cold start. On shared systems, measure queueing and abandoned requests; on an edge fleet, multiply update, replacement, backup, and travel time by node count, since operational labor usually outspends electricity. The common mistake: treating every saved minute as a full billing rate, when most converts to nothing anyone would invoice.
Payback month is the first month cumulative discounted benefit clears cumulative cost. ROI is net discounted benefit over discounted cost across a conservative horizon, residual value assuming the card goes economically obsolete before it physically dies. Rerun the model with productivity value at zero: clearing the bar only on speculative savings is the same red flag that makes a flat subscription the wrong shape for spiky demand. Compare against renting too: a short, honestly priced benchmark, storage and transfer costs counted, tells you more than another spec sheet.
State the break-even in checkable terms: accepted tasks per month, productive GPU hours, the maximum API cost per task that still loses to owning the hardware, minimum useful life. Set a retest date, plus earlier triggers: volume moves, a new model resets the quality bar, prices shift, the power contract changes, a component hits end of support.
I still don't have a clean answer for the risk premium itself. Everyone folds "no warranty, unknown wear, thin resale market" into a conservative haircut on useful life and calls it caution, but I've never seen that number be anything but a guess wearing a lab coat. I wouldn't run a fleet on it without opening the case more often than the spreadsheet says you need to.