AI hardware ROI for an edge or SBC AI fleet: valuing productivity without inventing savings
Every node in this kind of fleet runs its own power budget, and that budget doesn't care how good the model looks in a demo.
Hardware: Pi or Orange Pi boards, mini PCs, cameras, accelerators, storage, power, one management node, running sensing, OCR, vision, voice, or smart-home control near the data, against a cloud pipeline or plain deterministic rules with no model. The payoff: low power per node, no network hop, less data leaving, the case for local inference generally.
What matters is one reliable decision, made at the right place and time; tokens and accelerator-hours are just plumbing. A productivity claim needs a measured, verified completion time and a named economic use for the freed time. Keep prices live in a worksheet: they move faster than the writeup.
the ledger has to price what's actually incremental
Roughly: accepted tasks times minutes saved times loaded hourly value, divided by sixty, times a realization rate. Time the work with and without the system, include review, and don't round the rate up when the number disappoints.
Fixed: hardware, install, reserved capacity. Variable: electricity, tokens, transfer, support. Review swings both ways: one dedicated reviewer, fixed; per-exception correction, variable. Bill every incremental line minus what the org would buy anyway: a laptop isn't an AI cost, a memory upgrade to fit the model is.
Track the same fields across local, rented, and API routes:
- accepted tasks and severe failures
- input, cached input, reasoning, output
- retries, abstentions, human repair minutes
- queue, first response, completion, p95
- active, idle-resident, asleep, unavailable hours
- wall energy, cooling share, maintenance and incident time
Adjust for quality: two tries and five minutes of cleanup make a 'cheap' route not cheap, the cost belongs to the finished result. A narrow local model clearing one job reliably doesn't need pricing against a frontier API at full effort.
the constraint that breaks fleets is time, not silicon
The real risk is maintenance times node count: failed storage, install labor per unit, accelerators nobody uses consistently, boards stuck in prototype status nobody revisits. Build low, base, and high demand curves, a machine can't run at full utilization and still absorb an interactive request.
Meter idle states: sleeping overnight differs economically from a server holding models resident for instant answers. On a fleet, multiply update, replacement, backup, and travel time by node count, labor can outrun electricity as a cost.
Common mistake: billing every saved minute at full rate when nobody buys or avoids that time. Keep assumptions in a small visible table, date the volatile ones, skip the decimal precision that disguises a guess.
the number that matters is a boundary, not a verdict
Payback month is the first month cumulative discounted benefit clears cumulative cost; ROI is net discounted benefit over discounted cost. Both need a conservative useful life and honest residual value: hardware can keep running after the model catalog outgrows it or support ends.
Productivity needs its own realization factor: ten minutes saved doesn't mean ten minutes of billable output. Count it only when it avoids a hire, cuts spend, or shortens a binding queue. Run the model again at zero productivity; if it survives only on projected savings, look harder before signing.
Weigh the purchase against renting too, not just the API: hourly GPU rental earns its keep while demand and hardware shape stay uncertain, but count storage, transfer, and the instance somebody forgot to shut down.
The only durable output is a break-even boundary: accepted tasks per month, productive GPU hours, maximum API cost per task, minimum useful life, plus a retest date and its triggers: volume shifts, a model change, a price move, end of support on a core part.
What I haven't worked out is the half-finished fleet: a dozen nodes running for real, a handful more in a drawer because the pilot stalled. The spreadsheet calls that sunk cost either way, and calling it 'reserved capacity' just feels like the arithmetic this exercise exists to catch. No clean rule yet.