On-Premise, Again: What the Data Warehouse Years Taught Me About Local AI
Small business owners have started asking me whether they should run AI on their own machines instead of sending everything to ChatGPT, Copilot or Gemini. I understand why they ask, and I recognise the question, because I spent 2000 to 2002 on the other side of it.
Back then I was deploying SAP BusinessObjects and data warehouses for enterprise clients and training their teams, including assignments for the European Commission and recurring work in France and the UK. Everything we touched ran inside the client’s building. There was no cloud to send it to, so the debate about where data should live didn’t exist yet. You bought servers, a database licence, an ETL layer and the reporting tool, and then you hired or trained the people who could keep all of it running. A lot of those projects ended with a well-built warehouse that almost nobody opened, usually because the reports didn’t match how people actually worked, so they went back to their spreadsheets.
The companies that did get it right ended up owning something that took years to value properly: one controlled version of their own data, where they knew who could see what and what it cost to run. When software as a service arrived a decade later and removed the servers, the upgrades and the licence negotiations, most small businesses were right to take the deal, and for most of their software they still are. What got lost along the way was the habit of asking where the data goes once you press send.
I also watched, from inside, what happens to vendors. I went through two acquisitions and the dot-com bust, and the BI function I worked in was absorbed and eventually disappeared. The terms a provider offers you today are the terms that suit its business today, and I haven’t seen those hold for a full decade.
So with the same question back in 2026, I look at what has actually changed.
Open-weight models from the Llama, Mistral, Qwen and Gemma families are now good enough to classify documents, pull fields out of invoices, summarise internal notes or answer questions over a folder of company files, and they run on one workstation with a decent GPU. Two years ago that wasn’t realistic for a company of thirty people.
The data has changed too. People paste contracts, HR cases, client files and source code into chat windows, which is far more sensitive than anything they ever typed into a CRM. Samsung restricted generative AI tools in 2023 after engineers pasted confidential source code into ChatGPT, and Samsung has a security team to notice when that happens. Most SMBs would find out much later, if at all.
Regulation has moved as well. The GDPR was already there. The EU AI Act adds obligations in phases, and one of them, an AI literacy duty for companies that deploy AI, has applied since February 2025. The US CLOUD Act, which lets US authorities require US providers to hand over data they hold even when it’s stored in Europe, has gone from a legal footnote to something clients bring up in a first meeting.
The list of things that sink in-house projects is exactly the same as it was. Maintenance is the real cost, and it never appears on the hardware quote. The setup often depends on one enthusiastic person, and the day that person leaves it becomes a system nobody dares to touch. People bypass tools that don’t fit their work; with BI they went back to Excel, with AI they’ll go back to the free chatbot on their phone, which leaves you worse off than before. And “it’s on our own server” gets treated as a security policy, when a badly configured internal machine can be far easier to break into than a hyperscaler’s data centre.
What is genuinely better than in 2000 is that you no longer have to choose all or nothing. A warehouse was on-prem, full stop. Today you can run a small local model for the sensitive, repetitive tasks, send reasoning-heavy work on low-sensitivity data to a cloud model, and use an EU-hosted service for whatever sits in between.
That’s why I don’t treat local AI as a position to defend. It’s a deployment decision, and I’d make it on two criteria: how sensitive the data is, and how narrow and frequent the task is. Classifying client files, transcribing internal meetings or searching HR documents usually points to local. Drafting marketing copy or researching public information points to a good cloud model, which will beat anything you can run in the back office. Most of the rest fits an enterprise or EU-hosted plan with clear no-training terms.
Before arguing about where a use case runs, check that it deserves to run at all, which is what Screen It, the first part of my AI ROI Playbook, is for. The policy side, who may use which tool with which data, is in the Governance and Compliance series. The next two articles cover the data map and the real cost of running AI yourself.
The question I’d put to any owner considering a local setup is the one the BI years taught me to ask first: who will still be maintaining it in three years?