AI hardware ROI for an Apple Silicon local-model system: comparing the complete purchase price
Say you've settled on a number for the Mac Studio and a number for the rented A100 hour, and you feel good about the comparison. Stop. You're pricing a whole system against a sliver of one, and that gap is where the bad decisions happen.
The machine is Apple Silicon, mini, Studio, or MacBook, sized in memory and storage for the models you want plus the work you already do. The workload is quiet local development, document analysis, MLX experiments, and the occasional large model that needs a shared memory pool. The alternative is a discrete-GPU box plus an ordinary laptop, or a hosted API and no purchase, against one quiet, low-idle machine. The unit being priced is a productive local-model month, non-AI use included, not tokens per second or GPU-hours.
the box price is missing half the invoice
Capital cost is every component, tax line, shipping, spare, and displaced hardware needed to run the workload, priced from current quotes, power, and API rates in a worksheet you'll update, not prose you'll forget: prices drift, the shape doesn't.
installed_capex = compute + host + memory + storage + power + cooling + network + tax + setup
Build a bill of materials against that formula, then strip anything you'd buy anyway, a laptop for ordinary work isn't an AI cost, a memory upgrade to fit a bigger model is. Split fixed from variable: purchase and install on one side, electricity and paid calls on the other. Human review sits in between, a dedicated reviewer is fixed, spot-checking scales with how often it's wrong.
clock the workload before you clock the chip
Run the same tasks through local, rented, and API options and log what happens: accepted output against failures, the token mix, retries and correction minutes, first-token and p95 latency, hours active versus idle versus asleep. It's close to what a local-first-cascade setup already measures.
Quality adjustment isn't optional. A cheap model needing two attempts and five minutes of fixing has a cost belonging to the finished result, not the first wrong pass. A narrow local model that nails one job shouldn't be priced against a frontier API at full reasoning for that job.
unified memory doesn't forgive a wrong guess
The risk here is memory you can't upgrade later, steep price jumps between capacity tiers, workload-specific runtime support, and thin resale on odd configurations. I wouldn't bother modeling resale on an oversized config, nobody's buying it used in three years. Build low, base, and high demand curves including the daily peak and p95 latency: a maxed-out machine can't also absorb an interactive request.
Meter idle states honestly, a sleeping laptop differs economically from a box holding models resident, and on shared systems watch queueing and abandoned requests. Across a fleet, multiply update and replacement time by node count, labor beats the power bill. The recurring mistake: pricing a bare used GPU against a finished computer, or a turnkey API endpoint.
the payback date you can actually defend
Payback month is the first month cumulative discounted benefit clears cumulative cost. ROI over your horizon is net discounted benefit divided by discounted cost, both needing a conservative useful life and residual value, since hardware often outlives the models it was bought for.
Productivity gains need a realization factor first. Saving ten minutes a day doesn't manufacture ten minutes of billable output; it counts only when it avoids a hire, cuts spend sent elsewhere, or shortens a binding queue. Run it again with that term at zero: a purchase that only clears the bar on hoped-for time savings deserves suspicion, not a green light.
rent first, then write down when you'll check again
Rent and API routes deserve the same treatment: hourly GPU rental earns its keep while demand or hardware shape is uncertain, so fold in storage, image prep, transfer, minimum billing, and the forgotten idle instance, the same failure a subscription-wrong-for-ai commitment has. A short rental trial first is cheap insurance against a wrong purchase.
What survives is incremental installed cost attributable to this workload, nothing else. State break-even as accepted tasks per month, productive GPU-hours, max API cost per task, or a minimum useful life, checkable against reality, not a one-time story. Set a retest date and triggers for an earlier one: volume moves materially, a model resets the quality bar, API prices shift, or a component falls out of support.
The rule I keep: price the month, not the machine.