Infrastructure
A local AI server for your business: what it can really do and where it falls short
Running an AI model on your own server is easier and cheaper than it was two years ago, and for many tasks it is enough. But it isn't ChatGPT or Claude in your office. We cover the hardware you need, what it is good for, where the gap with frontier models shows and when it pays off.
Amud team · Reviewed on October 7, 2026 · 7 min read
What a local AI server is
It's a machine (a powerful computer, a workstation or a server in an office rack) that runs an open-weight language model: a model its creator publishes so anyone can download it and run it on their own hardware. The best-known families are Llama (Meta), Qwen (Alibaba), Mistral, Gemma (Google), DeepSeek and gpt-oss (OpenAI).
On that machine you install software that serves the model (Ollama, LM Studio, llama.cpp or vLLM, depending on the case), and from then on it's used like any other AI: from an internal chat, from an automation tool such as n8n or from the company's applications.
The difference from using ChatGPT or Claude is that the model and the data stay on your network. Nobody else processes what you send it.
What hardware you need
Memory is what matters. Compressed to 4 bits (the usual technique for running it locally, called quantisation), a model takes roughly half its parameter count in gigabytes, plus headroom for the text it processes. A 70-billion-parameter model needs about 40 GB just to load.
| Machine with one consumer GPU | Workstation | Server with several data-centre GPUs | |
|---|---|---|---|
| Available memory | 16–24 GB of video memory | 64–128 GB of video or unified memory | Hundreds of GB |
| Models it runs | Up to about 30 billion parameters | 70 to 120 billion | The largest open models |
| Good for | One person or a small team, well-defined tasks | A department, most office uses | Many concurrent users, complex agents |
| Indicative investment | €2,000–4,000 | €6,000–15,000 | Tens of thousands of euros or more |
Prices are indicative and move a lot with the memory and graphics card market. Two important caveats:
- Speed and concurrency. A model fitting in memory doesn't mean it answers quickly, or that it serves ten people at once. Each concurrent user and each page of context uses extra memory.
- Long context. Analysing an 80-page contract takes far more memory than a short question. Locally, context is expensive.
What it does well
In our experience, a mid-sized local model comfortably handles tasks that are well-defined, repetitive and easy to check:
- Classifying emails, tickets or documents by type, urgency or recipient.
- Extracting data from invoices, delivery notes, forms or contracts into a structured format.
- Summarising meetings, case files or long email threads.
- Searching internal documentation (known as RAG): answering questions from the company's manuals, procedures or records, citing the source document.
- Drafting replies, records or reports from a template.
- Transcribing audio from calls or meetings with models such as Whisper, which work very well locally.
- Assisting with code on simple tasks: explaining, documenting, writing tests.
Inside an automation workflow, a local model can take on the "smart" part (understanding an email, deciding who it goes to, pulling out the data) while programmed rules do the rest.
Where it falls short of OpenAI and Anthropic
This is where it pays to be realistic. The best open models are close to the frontier (according to Epoch AI, three to four months behind on average), but those leading open models have hundreds of billions of parameters, or even more than a trillion, and need data-centre servers. What fits on a consumer GPU takes between six months and a year to reach what frontier models already did. And what you'll run in the office is usually a compressed version that loses some quality.
In practice, the gap shows in:
Multi-step reasoning
A problem that requires chaining ten reasoning steps (reviewing a case file, cross-checking it against regulations and spotting an inconsistency) is where OpenAI and Anthropic models pull furthest ahead. A local model tends to get the simple steps right and lose the thread on long ones, or accept a conclusion without checking it.
Agents that use tools
An agent that decides what to look up, in what order, and corrects course when something fails needs a model that follows instructions very reliably over many iterations. Frontier models are trained specifically for this; mid-sized local ones drift sooner and more often.
Long documents
Commercial models work with hundreds of thousands of words of context at no hardware cost to you. Locally, every extra page costs memory and speed, and quality drops sooner when the text is long.
Edge cases
On the usual, the difference can be small. On the unusual (a badly scanned document, an ambiguous request, an exception to the rule) the frontier model gets it right more often. And in a business process, edge cases are the ones that cost money.
Maintenance
With a provider, the model improves on its own. Locally, updating the model, testing that the new one doesn't break what worked and keeping the server running is your job.
What it brings to confidentiality
It's the main reason businesses look at local AI, and it makes sense: if the model runs on your server, the data never leaves your network. No third party processes, stores or logs it, and there are no international transfers to justify. For highly sensitive material (medical records, legal case files, clients' financial data) it greatly simplifies the analysis.
But it isn't a complete answer:
- The server needs protecting too. Access control, a record of who queries what, backups, encryption and updates. A badly configured AI server on the internal network is another leak point.
- The GDPR still applies. Processing personal data with local AI still needs a legal basis, information for the people affected and data minimisation, even without an external provider.
- It isn't the only route. OpenAI, Anthropic and the large cloud providers offer business terms (no training on your data, limited retention, data processing agreements, processing in Europe) that in many cases let you use the best models with reasonable safeguards.
How to assess those providers, and what to require of them when confidentiality is critical for your business, deserves its own article, and we'll dedicate one to it. Meanwhile, AI agents in business: risks, GDPR and the AI Act covers the basic obligations.
Does it pay off?
A local server has a fixed cost: the hardware, the electricity (a GPU workstation running all day shows up on the bill) and above all the time of whoever installs, maintains and updates it. An external provider charges per use.
- With moderate use (a handful of people, a few hundred documents a month), paying per use is usually cheaper and gives access to better models.
- With high, steady volume of simple tasks (thousands of emails or documents a day), a local server can pay for itself in one or two years.
- When the data can't leave the company because of a contract or internal policy, the question stops being financial.
The approach that works best: hybrid
What we usually recommend isn't choosing one or the other but splitting the work:
- 01
Sensitive and simple work, locally
Classifying, extracting data or anonymising documents with personal data happens on your server, with nothing leaving it.
- 02
Complex work, with a frontier model
Hard reasoning goes to a provider on business terms, ideally with the data already anonymised in the previous step.
- 03
Rules decide what goes where
The workflow picks the route, not the user: by document type, by client or by whether it contains personal data.
- 04
Human oversight where it matters
Whatever the model, anything sent to a client or that changes data goes through a person.
That way you get the best of each side: privacy and a fixed cost for the volume, and the capability of the best models for what really needs it.
Before you buy the server
- Which specific tasks will it do, and how many a day?
- Have those tasks been tested with an open model of the size you'll be able to run, on real documents?
- How many people will use it at once?
- Who will maintain and update it?
- Is there data that, by contract or policy, can't leave the company?
- Which tasks will still need a frontier model?
If you want to work out which part of your processes a local model can handle and which needs more, AI agents for business explains how we approach it.
Sources
Frequently asked questions
Is a local model as good as ChatGPT or Claude?
Not quite. The best open models trail the frontier by a few months, but those models need servers with several data-centre GPUs. What fits on an office computer or server performs roughly like the commercial models of one or two years ago: very well on well-defined tasks, worse on long reasoning and on agents that chain many steps.
What computer do I need to run AI locally?
It depends on the model size. A machine with a 16 to 24 GB graphics card runs small and mid-sized models, enough to classify, summarise or extract data. Models of 70 to 120 billion parameters need a workstation with 64 to 128 GB of video or unified memory. The largest models require data-centre servers.
Does running AI locally guarantee confidentiality?
It helps a lot, because the data never leaves your network, but it doesn't guarantee it on its own: the server needs access control, logging, backups and updates like any other system. And it isn't the only route: the large providers offer business terms worth knowing.
Is a local server cheaper than paying OpenAI or Anthropic per use?
Only with high, steady volume. On top of the hardware come electricity, maintenance and the time of whoever runs it. With moderate use, paying a provider per use is usually cheaper and gives access to better models.
Keep reading
Blog
AI agents in business: risks, human oversight, GDPR and the AI Act
An AI agent that can act in your systems is very useful and, without controls, risky. These are the risks we see in practice, the controls that neutralise them and what the GDPR and the EU Artificial Intelligence Act say.
Service
AI agents for businesses
We build artificial intelligence agents that do real work inside your company: answering clients, reading and sorting documents, qualifying leads or preparing case files. With clear limits and a person approving anything important.
Blog
n8n vs Make vs Zapier: which one to automate your business
All three tools connect apps and automate processes, but they charge differently, are hosted in different places and fit different projects. This is the comparison we make before choosing.
Which task would you like off your plate?
Tell us how your team works. In a free 30-minute session we'll tell you what can be automated, how much you'd save and what isn't worth it.
Book a free meeting