Infrastructure

A local AI server for your business: what it can really do and where it falls short

Running an AI model on your own server is easier and cheaper than it was two years ago, and for many tasks it is enough. But it isn't ChatGPT or Claude in your office. We cover the hardware you need, what it is good for, where the gap with frontier models shows and when it pays off.

Amud team · Reviewed on October 7, 2026 · 7 min read

What a local AI server is

It's a machine (a powerful computer, a workstation or a server in an office rack) that runs an open-weight language model: a model its creator publishes so anyone can download it and run it on their own hardware. The best-known families are Llama (Meta), Qwen (Alibaba), Mistral, Gemma (Google), DeepSeek and gpt-oss (OpenAI).

On that machine you install software that serves the model (Ollama, LM Studio, llama.cpp or vLLM, depending on the case), and from then on it's used like any other AI: from an internal chat, from an automation tool such as n8n or from the company's applications.

The difference from using ChatGPT or Claude is that the model and the data stay on your network. Nobody else processes what you send it.

What hardware you need

Memory is what matters. Compressed to 4 bits (the usual technique for running it locally, called quantisation), a model takes roughly half its parameter count in gigabytes, plus headroom for the text it processes. A 70-billion-parameter model needs about 40 GB just to load.

Machine with one consumer GPUWorkstationServer with several data-centre GPUs
Available memory16–24 GB of video memory64–128 GB of video or unified memoryHundreds of GB
Models it runsUp to about 30 billion parameters70 to 120 billionThe largest open models
Good forOne person or a small team, well-defined tasksA department, most office usesMany concurrent users, complex agents
Indicative investment€2,000–4,000€6,000–15,000Tens of thousands of euros or more

Prices are indicative and move a lot with the memory and graphics card market. Two important caveats:

  • Speed and concurrency. A model fitting in memory doesn't mean it answers quickly, or that it serves ten people at once. Each concurrent user and each page of context uses extra memory.
  • Long context. Analysing an 80-page contract takes far more memory than a short question. Locally, context is expensive.

What it does well

In our experience, a mid-sized local model comfortably handles tasks that are well-defined, repetitive and easy to check:

  • Classifying emails, tickets or documents by type, urgency or recipient.
  • Extracting data from invoices, delivery notes, forms or contracts into a structured format.
  • Summarising meetings, case files or long email threads.
  • Searching internal documentation (known as RAG): answering questions from the company's manuals, procedures or records, citing the source document.
  • Drafting replies, records or reports from a template.
  • Transcribing audio from calls or meetings with models such as Whisper, which work very well locally.
  • Assisting with code on simple tasks: explaining, documenting, writing tests.

Inside an automation workflow, a local model can take on the "smart" part (understanding an email, deciding who it goes to, pulling out the data) while programmed rules do the rest.

Where it falls short of OpenAI and Anthropic

This is where it pays to be realistic. The best open models are close to the frontier (according to Epoch AI, three to four months behind on average), but those leading open models have hundreds of billions of parameters, or even more than a trillion, and need data-centre servers. What fits on a consumer GPU takes between six months and a year to reach what frontier models already did. And what you'll run in the office is usually a compressed version that loses some quality.

In practice, the gap shows in:

Multi-step reasoning

A problem that requires chaining ten reasoning steps (reviewing a case file, cross-checking it against regulations and spotting an inconsistency) is where OpenAI and Anthropic models pull furthest ahead. A local model tends to get the simple steps right and lose the thread on long ones, or accept a conclusion without checking it.

Agents that use tools

An agent that decides what to look up, in what order, and corrects course when something fails needs a model that follows instructions very reliably over many iterations. Frontier models are trained specifically for this; mid-sized local ones drift sooner and more often.

Long documents

Commercial models work with hundreds of thousands of words of context at no hardware cost to you. Locally, every extra page costs memory and speed, and quality drops sooner when the text is long.

Edge cases

On the usual, the difference can be small. On the unusual (a badly scanned document, an ambiguous request, an exception to the rule) the frontier model gets it right more often. And in a business process, edge cases are the ones that cost money.

Maintenance

With a provider, the model improves on its own. Locally, updating the model, testing that the new one doesn't break what worked and keeping the server running is your job.

What it brings to confidentiality

It's the main reason businesses look at local AI, and it makes sense: if the model runs on your server, the data never leaves your network. No third party processes, stores or logs it, and there are no international transfers to justify. For highly sensitive material (medical records, legal case files, clients' financial data) it greatly simplifies the analysis.

But it isn't a complete answer:

  • The server needs protecting too. Access control, a record of who queries what, backups, encryption and updates. A badly configured AI server on the internal network is another leak point.
  • The GDPR still applies. Processing personal data with local AI still needs a legal basis, information for the people affected and data minimisation, even without an external provider.
  • It isn't the only route. OpenAI, Anthropic and the large cloud providers offer business terms (no training on your data, limited retention, data processing agreements, processing in Europe) that in many cases let you use the best models with reasonable safeguards.

How to assess those providers, and what to require of them when confidentiality is critical for your business, deserves its own article, and we'll dedicate one to it. Meanwhile, AI agents in business: risks, GDPR and the AI Act covers the basic obligations.

Does it pay off?

A local server has a fixed cost: the hardware, the electricity (a GPU workstation running all day shows up on the bill) and above all the time of whoever installs, maintains and updates it. An external provider charges per use.

  • With moderate use (a handful of people, a few hundred documents a month), paying per use is usually cheaper and gives access to better models.
  • With high, steady volume of simple tasks (thousands of emails or documents a day), a local server can pay for itself in one or two years.
  • When the data can't leave the company because of a contract or internal policy, the question stops being financial.

The approach that works best: hybrid

What we usually recommend isn't choosing one or the other but splitting the work:

  1. 01

    Sensitive and simple work, locally

    Classifying, extracting data or anonymising documents with personal data happens on your server, with nothing leaving it.

  2. 02

    Complex work, with a frontier model

    Hard reasoning goes to a provider on business terms, ideally with the data already anonymised in the previous step.

  3. 03

    Rules decide what goes where

    The workflow picks the route, not the user: by document type, by client or by whether it contains personal data.

  4. 04

    Human oversight where it matters

    Whatever the model, anything sent to a client or that changes data goes through a person.

That way you get the best of each side: privacy and a fixed cost for the volume, and the capability of the best models for what really needs it.

Before you buy the server

  • Which specific tasks will it do, and how many a day?
  • Have those tasks been tested with an open model of the size you'll be able to run, on real documents?
  • How many people will use it at once?
  • Who will maintain and update it?
  • Is there data that, by contract or policy, can't leave the company?
  • Which tasks will still need a frontier model?

If you want to work out which part of your processes a local model can handle and which needs more, AI agents for business explains how we approach it.

Sources

Frequently asked questions

Is a local model as good as ChatGPT or Claude?

Not quite. The best open models trail the frontier by a few months, but those models need servers with several data-centre GPUs. What fits on an office computer or server performs roughly like the commercial models of one or two years ago: very well on well-defined tasks, worse on long reasoning and on agents that chain many steps.

What computer do I need to run AI locally?

It depends on the model size. A machine with a 16 to 24 GB graphics card runs small and mid-sized models, enough to classify, summarise or extract data. Models of 70 to 120 billion parameters need a workstation with 64 to 128 GB of video or unified memory. The largest models require data-centre servers.

Does running AI locally guarantee confidentiality?

It helps a lot, because the data never leaves your network, but it doesn't guarantee it on its own: the server needs access control, logging, backups and updates like any other system. And it isn't the only route: the large providers offer business terms worth knowing.

Is a local server cheaper than paying OpenAI or Anthropic per use?

Only with high, steady volume. On top of the hardware come electricity, maintenance and the time of whoever runs it. With moderate use, paying a provider per use is usually cheaper and gives access to better models.

Keep reading

Which task would you like off your plate?

Tell us how your team works. In a free 30-minute session we'll tell you what can be automated, how much you'd save and what isn't worth it.

Book a free meeting