← all posts
// economics · hardware-roi

AI hardware ROI for an Apple Silicon local-model system: break-even against commercial APIs

A Mac Studio does not pay for itself just because you stopped paying an API bill.

Owners repeat that line to feel smart about the purchase; it's true only past a break-even point almost nobody calculates. Local wins once enough accepted work crosses the machine. Not before.

The machine here is unified-memory Apple Silicon, sized for target models plus normal apps: local development, document analysis, MLX work, an occasional model wanting a shared pool. The alternative: a discrete-GPU workstation plus a laptop, or a hosted API, the fork mapped in local-first cascade decisions. Apple Silicon buys a quiet, low-idle box and memory never split between VRAM and system RAM.

The unit that matters is a productive local-model month, counting the machine's ordinary non-AI use too; tokens and GPU-hours are only intermediate numbers. An API bill compared on input-token price alone is close to useless: you need quality, cached and output tokens, tool calls, retries, review cost. Keep prices in a live spreadsheet, not fixed prose; they move every quarter, the formula doesn't.

What actually gets counted

Break-even in tasks equals fixed local cost divided by the gap between accepted-task cost through the API and the marginal local cost of the same task, priced per accepted result. Purchase and reserved capacity are fixed; electricity, tokens, and transfer scale with volume; a dedicated reviewer is fixed, fixing exceptions scales with error rate. List every incremental line, then subtract what you'd buy anyway: the laptop you already needed isn't an AI cost, the memory bought only so a model fits, is.

The log nobody keeps

Log more than pass or fail across local, rented, and API routes: accepted tasks and failures, cached and output tokens, retries and repair minutes, first-token time, active versus idle hours. Skip the quality adjustment and the comparison lies: a cheap model needing two attempts and a five-minute fix isn't cheap, its cost belongs to the completed result. Don't price a narrow local specialist against a frontier API at max effort.

The real risks: non-upgradeable memory, ransom-priced capacity, one-framework lock-in, weak resale on an odd build. Sketch low, base, and high demand; full utilization can't also answer prompts promptly, and batch work fills the gaps only if genuinely needed. The common mistake: raw token counts compared across systems with different success rates.

Buy signal, not a feeling

Payback month is the first month cumulative discounted benefit clears cumulative cost. ROI over your horizon is net discounted benefit over discounted cost, both only as good as your estimate of useful life and residual value. A machine can still boot and still be economically dead once model growth or dropped support passes it by.

Productivity claims need a realization factor: saving ten developer minutes isn't revenue unless it avoids a hire, cuts outsourced spend, or raises output. Run the model once with that value at zero; distrust break-even that only happens on speculative time savings.

Compare against renting too: hourly GPU rental earns its keep while demand or hardware shape stays uncertain, and a short benchmark rental beats an expensive wrong purchase once storage, transfer, and minimum billing are counted. Buy when demand clears break-even inside your ownership horizon, stated checkably: accepted tasks per month, GPU hours, a cost ceiling per task. Set a retest date and recalculate when volume shifts, the quality bar changes, or prices move.

What I still do not have a clean answer for is resale. Nobody wants a used Mac with memory sized for one model generation, and I have no honest way to price that residual once the model that justified buying it goes obsolete. I price it at zero and move on. Probably wrong.

#hardware-roi#cost#self-hosting