AI hardware ROI for a personal AI workstation: electricity and cooling economics
Here's the rule: model the whole box in every power state it actually sits in, or don't bother modeling at all. Half a spreadsheet is worse than no spreadsheet; it hands you false confidence dressed up in decimals. People price a workstation off GPU wattage times generation time, forget the rest of the machine exists, and call it math.
This is a personal AI workstation: one developer's desktop, a GPU with headroom, RAM, NVMe, a power supply that won't trip a breaker, running interactive coding, private document work, experiments, and occasional batches. The alternative is a mix of API calls and whatever laptop is already on the desk, a local-first cascade question restated in dollars. The upside is low-latency, private access to open-weight models without asking permission per token, priced in months of verified developer work, not tokens or GPU-hours, just intermediate readings.
the wattage sticker lies by omission
Energy cost is wall kilowatts times active hours times electricity rate times a cooling multiplier, and that last term is the one people drop. Air conditioning a home office isn't free, and poor airflow pushes it up regardless. Meter the whole system across every state, sleep, idle, model-resident idle, full load, weighted by time actually spent there.
Split fixed from variable early. Accelerator, case, and install labor are fixed and sunk; electricity, API overflow, and support time move with usage; human review is fixed if dedicated, variable if not. Build a real cost architecture, not a one-line guess: price every incremental part, then subtract what you'd have bought anyway. The laptop already on your desk isn't an AI cost. A RAM upgrade bought solely so a 14B-class GGUF fits in memory is.
idle hours are where the estimate quietly dies
The bigger risk isn't the hardware bill: low utilization dressed up as savings, plus upgrading on every model release and calling the hobby a business expense. Build low, base, and high demand curves with real peak load and p95 latency: a machine at full tilt has nothing left for an interactive request. Batch work can soak up quiet hours, but only if that work needed doing anyway.
Meter idle honestly. A workstation that sleeps overnight shares nothing, economically, with a server holding a model resident all night for instant answers. Scale to a fleet and labor buries electricity.
Quality matters too, not just speed. A cheap model needing two attempts and five minutes of cleanup isn't cheap: the cost belongs to the finished result, not the first draft. Don't price a small local model against a frontier API at max effort for a job it clears first try, either; log dated sources for moving numbers.
payback is a boundary you set, not a headline you chase
Payback month is the first month cumulative discounted benefit clears cumulative cost. ROI over your horizon is net discounted benefit over discounted cost, and both need a conservative useful life and real residual value, since hardware often outlives the model growth or dropped support that made it pointless. Productivity savings need a realization factor: shaving ten minutes off a task doesn't hand you ten billable minutes unless it avoids a hire, cuts outsourced spend, or clears a real bottleneck. Run the model once with productivity value at zero, and if the purchase only pencils out once speculative time goes back in, look at it twice.
Price renting too. Hourly GPU rental earns its keep while demand shape is uncertain, if you count persistent storage, transfer, minimum billing, and the instance you forgot was running. A short rental benchmark beats an expensive wrong purchase.
Joules per accepted task is the only unit that doesn't lie to you six months later.
State your break-even in operational terms: accepted tasks per month, productive GPU hours, the max API price per task that flips the decision, plus a sleep floor for idle hours. That's revisable, not a narrative you defend. Set a retest date and triggers: volume moves, a model changes the quality gate, prices shift, the electricity contract renews, or a part hits end of support.
None of this spreadsheet discipline saves you from the real failure mode: building it once, feeling rigorous, and never opening it again until the GPU is the wrong shape for the job.