AI hardware ROI for a used-GPU inference build: the utilization curve
Second-hand accelerators come with no warranty, unknown wear, and a resale value that can evaporate the day a new architecture ships.
The build here is a used GPU or two, a case with a compatible motherboard, memory, a power supply sized for spikes, cooling, and a replacement plan, for steady local inference, where VRAM beats the newest tensor cores. The alternatives are new consumer cards, a refurbished workstation, or paying per token to someone else's cluster, a bet covered in why a subscription is the wrong frame for AI costs. None of it wins by default: a cheap rig can still lose to an API bill waiting all day for work that never comes.
The hour is the unit, not the token (hardwareroi)
What you're producing is a quality-gated inference hour: usable time spread across the hardware's remaining useful life. Tokens and GPU-hours are just readings on the way, tracked as accepted-vs-failed tasks, token type, retries, repair minutes, and latency through p95, across local, rented, and API routes. The formula is blunt: annual fixed cost divided by productive hours, plus whatever moves with the work. Fixed costs (purchase, install, reserved capacity) apply regardless of hours run; variable costs (power, tokens, transfer, support) track the work; human review is fixed if dedicated, variable if per-exception.
Price every incremental part: card, host, memory, storage, power, cooling, tax, installation, then subtract whatever the org would have bought anyway; a laptop the team already needed isn't an AI cost, a memory upgrade bought just to fit a model, is. Two attempts and five minutes of cleanup means the real cost sits with the finished result, not the first draft; a local model that reliably nails one task shouldn't be priced against a frontier API at max reasoning.
Wear, warranty, and the resale cliff
Buying secondhand moves risk more than cost: wear nobody measured, no warranty, an unbudgeted platform upgrade, higher power draw than the newer part, downtime chasing a replacement, and a resale market that can go quiet exactly when you need to sell.
Build three demand curves, low, base, high, including daily peaks and p95 latency: a machine scheduled at full utilization on paper still has to answer an interactive request the moment it lands. Count batch work only when the organization needs it.
Idle capacity still draws power
The rule underneath: fixed cost spreads across hours the hardware earns its keep, while idle time still draws on capital and, often, electricity. Run the numbers with current prices, in a worksheet you update, not frozen prose; prices move faster than the decision's shape.
Meter every hour, because utilization hides several states:
- queued and actively serving a request
- resident-idle, model loaded, nothing in flight
- asleep, powered down between shifts
- unavailable, down for maintenance or a failed part
A workstation that sleeps overnight has a different cost profile than a server holding models resident for instant answers. On shared systems, watch the queue and count abandoned requests; on an edge fleet, multiply update, replacement, backup, and site-visit time by node count, since labor can outrun power. The common mistake: dividing by the theoretical 8,760 hours in a year instead of hours you'll use, kept in a small dated table, uncertainty shown, not buried in decimal places.
The zero-productivity gut check
Payback month is the first month cumulative discounted benefit clears cumulative cost; ROI is net discounted benefit over discounted cost. Both lean on a conservative useful life and residual value, and a card can outlive its usefulness once the model landscape or its support moves on.
Productivity claims need a haircut: ten minutes saved per developer doesn't turn into ten minutes of revenue. Count it only when it avoids a hire, cuts outside spend, adds to shipped work, or clears a bottleneck. Run the model again at zero. If it only clears the bar on hoped-for savings, treat that number with suspicion.
Rent first, the same local-first cascade logic, while demand and hardware shape stay uncertain, and run a short benchmark against an expensive wrong purchase, pricing it in full: storage, image prep, transfer, minimum billing, spin-up automation, the idle instance left running.
Where the break-even line actually sits (hardwareroi)
Size the build for attainable utilization, not utilization that flatters the spreadsheet, while keeping the latency headroom users need. State the boundary in operational terms: accepted tasks per month, productive GPU hours, the max you'd pay an API per task before the math flips, or the minimum useful life needed. That boundary you revisit in five minutes next quarter, not argue over.
Set a retest date and the triggers that move it earlier: volume shifts enough to matter, a model changes the quality gate, API prices move, the electricity contract changes, or a component hits end of support. That payoff is a relationship, not a fixed property of the card: one workload, one alternative, dated assumptions, all temporary.
Given the choice, I round utilization down and eat the hit on cost per hour rather than round up and get blindsided in month four. The price for that caution is a build that looks underused on paper longer than I'd like to admit, and I'm fine paying it.