← Back to Articles

The Real Cost of Running AI in Your Back Room

The usual pitch for local AI is that you stop paying subscriptions. You do stop paying per user or per request, and in exchange you pay for hardware, electricity, setup and the hours of whoever keeps the thing running. Whether that’s worth it depends on your volumes and your data, and it’s much cheaper to work out on paper than after the hardware has arrived.

What I learned running it myself

I’ve been testing local models on my own desktop: an Nvidia card with 16 GB, 64 GB of RAM, LM Studio to run the models and OpenCode as the harness that connects them to my IDE. To be clear, this is definitely not a setup for an enterprise, and not for a company of 20 people either. A setup that stays affordable for a small company would start with a 32 GB graphics card. With the new mixture-of-experts models and quantization, that kind of configuration can run decent models for coding and for tasks of basic to moderate complexity, and it’s more than enough for simple writing help and the everyday tasks most employees need.

So far I’ve used the coding models to build simple test websites rather than on my real codebases. The code itself is correct. The design is not, but that looks like a configuration question, since the design side depends on the front-end frameworks the model has to work with, and I expect it can be solved.

What surprised me more is the time it takes, even for a personal test setup: research, configuring the local environment, finding out which open models exist and what they can do, connecting the harness to my IDE, understanding the logic behind the models and how to set their parameters. And new models come out regularly, with new ways of making better use of the same hardware, so part of that work never really stops. For a company, a realistic first deployment is more like ten days of work, because a lot has to happen before anything gets installed: analysing the company’s needs, choosing the use cases, mapping which data is involved, sizing the hardware and picking the models, then integrating, testing on real examples and training the people who will use it.

What you pay for

I’m deliberately not putting prices on the hardware here. The right configuration depends on how many people use it, for which tasks and with which models, and that needs a proper analysis for each company. What I can give you is the list of lines to price, and where the surprises are.

Cost line What drives it How to estimate it
Hardware GPU memory above all; 32 GB is a realistic starting point for a small company Quotes for two or three configurations, sized on your actual use cases
Setup and integration Connecting the models to the tools people already use Around ten days for a first deployment, needs analysis included
Electricity Whether the machine runs all day or on demand Meter it for a month instead of estimating
Maintenance and learning New models, updates, parameter tuning, fixing what breaks Hours per month, with a name next to them

The last line is the one people leave out. Setup is a project with a start and an end, but keeping up with new models, updates and tuning is recurring work, and in a company those hours are someone’s working time taken from something else, or an outside partner’s invoice.

What a local model can and can’t do

Task Local model on an affordable setup Why
Classifying documents or emails Strong fit Narrow, repetitive, easy to test
Extracting fields from invoices, forms, contracts Strong fit Structured output, clear success criteria
Transcribing internal meetings Strong fit Mature open-source speech models exist
Simple writing help and everyday employee tasks Good fit Enough for most daily needs on a 32 GB card
Searching and answering questions over internal documents Good fit Quality depends more on your document setup than on the model
Coding Good fit Correct code in my tests; design and front-end work need configuration
Polished client-facing text Weak fit Noticeable quality gap vs. top cloud models
Complex analysis, multi-step reasoning, strategy Weak fit This is where frontier cloud models earn their price

If your use case sits in the last two rows, going local saves money at the expense of quality. Ask the people who will use it whether they accept that, because if they don’t, they’ll quietly go back to the chatbot on their phone and you’ll end up paying for both.

Break-even

The comparison is what the same task would cost you in the cloud each month against the full monthly cost of running it yourself.

Local pays off when: cloud cost > (H + S) ÷ 36 + E + M — H is hardware, S is setup, E is monthly electricity and M is monthly maintenance, over a 36-month horizon. Put your own quotes and your own hours into it. If the maintenance term alone is bigger than what you currently pay for the cloud tools you’d replace, you already have your answer.

In practice three patterns come up. With a high volume on a narrow task, say thousands of documents classified or extracted every month, cloud costs grow with each request while the local cost stays flat, and local usually wins. When a few people use AI for varied writing, research and analysis, cloud spend stays modest and quality matters more, so the cloud usually wins by a wide margin. And when the data map from the previous article says a task can’t go to a public cloud at all, cost stops being the deciding factor; the real comparison becomes local against an EU-hosted service, and the same formula lets you make it honestly.

Most SMBs end up running both: one local setup for sensitive, high-volume work, and business-tier or EU-hosted tools for everything else.

Before you buy hardware

Where this fits

This is one decision inside the process my AI ROI Playbook covers. Screen It checks whether a use case deserves attention at all. Do the Numbers gives the full ROI calculation, of which this break-even is one input. Get It Approved turns it into a business case your partners or board will sign off, and Measure What Happened is where you check, six months in, whether the local setup delivered what the business case promised.

If you only do one thing from this article, fill in the maintenance line with a name and a number of hours. And budget the first deployment in days, around ten for a small company, a good part of them spent before anything is installed.