AI hardware ROI for an edge or SBC AI fleet: electricity and cooling economics
The number on your accelerator's spec sheet is nearly useless for costing a fleet. It's not wrong, just incomplete: no host board, no peripherals, no power-supply losses, no idle hours.
Picture Raspberry Pis, Orange Pis, mini PCs, cameras, accelerators, storage, enclosures, one management node, running always-on sensing, OCR, vision, voice, or narrow automation next to the data, instead of a cloud pipeline, a central server, or plain sensors: lower per-node draw, faster response, less data leaving the building, placement near the sensor. What you're buying is one reliable decision made where it matters; tokens and accelerator-hours are proxies.
Model all of it, plus duration and cooling overhead, or you're guessing. Keep rates and quotes in a worksheet, not in prose. Prices drift. The decision boundary outlives them.
Count everything the accelerator drags in with it
Formula: wall kilowatts times active hours times electricity rate times a cooling factor. Meter the whole system, not the chip, across sleep, idle, model-resident, and active states, at the duty cycle you observe. Split fixed from variable: hardware and installation are fixed; electricity, tokens, transfer, and support scale with use.
List every incremental component, accelerator through installation, minus anything you'd buy anyway: a laptop your engineer owns isn't an AI cost, a memory upgrade bought only to fit a bigger model is. Run the workload through local, rented, and API routes, tracking failures, repair minutes, and p95 latency. A model needing two attempts and a five-minute fix isn't cheaper; cost belongs to the finished result.
Multiply the maintenance, not just the milliwatts
The real risk is rarely the power draw. It's maintenance scaling with node count, storage failing at the worst time, installation labor, fragmented accelerators, and boards stuck at prototype status. Build low, base, and high demand curves, watch daily peaks and p95 latency: a box at full utilization has no headroom for an interactive request. Batch work counts only if it genuinely exists.
A workstation sleeping overnight has different economics than a server holding models resident for instant response. Across a fleet, multiply update, replacement, backup, and travel time by node count; labor can dwarf electricity before anyone notices. The common mistake is multiplying GPU TDP by generation time while ignoring idle and the rest of the machine. Keep assumptions in a short table, date what's volatile, and don't hide uncertainty in decimals.
Payback is a date, not a vibe
Payback month is the first month cumulative discounted benefit passes cumulative cost; ROI over your horizon is net discounted benefit over discounted cost. Both need a conservative useful life and residual value: hardware often outlives its economic usefulness before it breaks. Productivity gains need a realization factor: saving ten minutes doesn't hand you ten billable minutes, it counts only if it avoids a hire, cuts outsourced spend, or shortens a genuinely binding queue. Run the model once with that value at zero; if the hardware clears the bar only on hoped-for savings, be suspicious.
Compare against renting: hourly GPU rental earns its keep while demand and hardware shape are uncertain, the same case against a flat subscription for API-priced work. Include storage, transfer, minimum billing, and the idle instance someone forgot to shut down; a short benchmark beats an expensive wrong purchase.
The durable move: optimize joules per accepted task, schedule sleep once availability stops justifying idle draw, the same local-first cascade reasoning applied to hardware. State the break-even boundary: accepted tasks per month, productive GPU hours, a ceiling on API cost per task, or a minimum useful life. Set a retest date with triggers: volume shifts, a model changes the quality gate, prices move, or a component reaches end of support.
None of this is hard, just tedious, and for a small pilot fleet I skip the spreadsheet and watch the power meter for a week, exactly the shortcut this piece told you not to trust.