AI hardware ROI for a shared team GPU server: comparing the complete purchase price
Price the fully installed system against the fully loaded cost of the alternative, at your own utilization, or don't run the numbers at all. That's the whole rule. Everything below just proves it holds for a shared team GPU box, and points at where people quietly cheat on it.
Picture the machine: a multi-user Linux host, one or more accelerators, server memory, redundant storage, networking, monitoring, someone keeping it alive, feeding coding assistants, retrieval, extraction, evaluation, and scheduled batch work for several engineers. The competing option is commercial APIs or a rented endpoint sized for the same traffic. What you're buying is pooled utilization, data staying where you put it, model versions that hold still, and capacity you own instead of rent hourly. Price it against one unit, an accepted task at the latency your team needs, not the tokens along the way, same discipline as the cost-architecture teardown: price the system, not the sticker.
What actually belongs on the invoice
Installed capital cost is compute plus host plus memory plus storage plus power and cooling plus network plus tax plus setup, full stop; miss one and it's fiction dressed as math. Build one complete bill of materials per route, then subtract whatever the org would buy anyway: a developer's laptop stays out of the AI column, a memory stick bought only because a model needs it doesn't. Split what's left into fixed and variable: purchase, install, and reserved capacity are fixed, electricity, metered tokens, transfer, and support move with the work, and human review sits in between, a dedicated reviewer is fixed, one who shows up only on failure is variable. Run the workload through every route, local, rented, API, and record accepted tasks against failures, the token mix, retries and human-fix minutes, latency to first token and completion at your promised percentile, and hours spent active, idle-resident, asleep, or down. Refresh those prices in a worksheet, not here; a quote from months back is a placeholder, not a fact.
Idle time is where the spreadsheet lies
Utilization doesn't get to be an assumption, you model it. The real risk isn't the accelerator dying: it's queues piling up, capacity idle, someone's afternoon spent keeping the thing alive, a single host taking the team down when it hiccups, demand outgrowing the shape you sized for. Build a low, an expected, and a high demand curve that respects your daily peak and promised p95, since you can't run full occupancy and still absorb an interactive burst. Batch jobs soak up slack between peaks, but only if that work was happening anyway.
A bare used GPU on a bench is not the same purchase as a finished, powered, monitored box, and comparing the two is the most common way this whole exercise gets rigged.
Meter idle states honestly: a workstation sleeping overnight differs from a server holding several models resident so nothing waits on a cold load. On a shared box, track queueing and abandoned requests, that number says more about real capacity than any spec sheet. An edge fleet changes the math: multiply update, replacement, backup, and travel time by node count, operational labor there routinely outweighs the power bill. Keep assumptions in a small table with sources and dates, and skip fake precision, it just hides a round number as four decimal places.
A retry doesn't count as fast
Quality adjustment isn't optional. If the cheap route needs two attempts and five minutes of someone fixing the output, that five minutes is part of its price, not a footnote. You're paying for the completed result, not the first draft. Reverse it too: a small local specialist that reliably nails one narrow task shouldn't get judged against a frontier API at maximum reasoning effort, you were buying the task done right, not reasoning effort. Same discipline applies to productivity, where the temptation to inflate runs strongest. Saving a developer ten minutes doesn't hand you ten minutes of revenue; it counts only when it avoids a hire, cuts spend sent outside, or shortens a queue that was blocking someone. Run the model again with that term at zero. If it only clears break-even on assumed time savings, that's the finding.
Break-even is a boundary, not a story
Payback month is the first month cumulative discounted benefit clears cumulative cost. ROI over your horizon is net discounted benefit divided by discounted cost, both only as honest as the useful-life and residual-value guesses underneath. I wouldn't spend much effort pinning residual value to the dollar; nobody resells a used inference box for what the spreadsheet claims, and hardware usually goes obsolete before it physically dies. Price hourly rental too, not just the API: a rented accelerator earns its keep while demand and hardware shape are still uncertain, counting persistent storage, image setup, transfer, minimum billing, and the instance somebody forgot to shut down. If your default is a subscription because it's easier to expense, that instinct is worth checking. A short rental run is cheap insurance against an expensive wrong purchase. What survives is the incremental installed cost of the workload, plus a break-even line in operational terms: accepted tasks per month, productive accelerator-hours, the max API cost per task you'll tolerate, or the minimum useful life the hardware needs to hit. Put a retest date on the calendar, and name what moves it early: real volume change, a model resetting the quality bar, an API price move, an electricity contract change, a component falling out of support.
The one rule I'd keep: ROI is a relationship between one workload, one alternative, and a set of assumptions with a date on them, never a property stamped onto the GPU itself.