All articles

Running a Private LLM with Ollama: A Practical Guide for Companies

How to run open-weight language models on your own servers with Ollama: when self-hosting makes sense, sizing GPU memory, Ollama versus vLLM, securing the API, and evaluating quality on your own data.

AI Development|Published |10 min read
Server racks in a data centre hosting a private language model

A private LLM is a language model that runs on infrastructure the company controls, so prompts, documents, and answers never leave its network. Open-weight models such as Gemma, Qwen, Llama, and DeepSeek have made this practical, and Ollama has made it simple to start. Getting a demo running takes an afternoon. Making it something the business uses every day takes longer, and this guide covers the decisions that come up along the way.

When a Private LLM Makes Sense

Self-hosting is not automatically the better option. Managed APIs such as Claude, GPT, and Gemini offer stronger reasoning, no hardware to buy, and a fast path to a proof of concept. A private LLM earns its place when the data cannot leave the organisation, when request volume is high enough that per-token pricing becomes the dominant cost, or when the system must keep working without an internet connection.

FactorSelf-hosted LLMManaged API
Data locationStays on servers you controlSent to the provider, usually under zero-retention terms
Reasoning qualityGood for focused tasks such as extraction, classification, and grounded Q&AStrongest available for complex, open-ended reasoning
Cost modelUpfront GPU investment or rental, flat cost per monthPay per token, grows with usage
OperationsYour team runs, patches, and monitors the serviceProvider handles capacity and uptime
Best fitRegulated data, internal documents, steady high volumePrototypes, low volume, tasks needing frontier-grade answers

Many production systems end up hybrid. Sensitive documents are processed by a self-hosted model, while non-sensitive tasks that benefit from stronger reasoning are routed to an API. Designing that routing early is easier than retrofitting it after the first model has been integrated everywhere.

Sizing GPU Memory for an Open-Weight Model

The first hardware question is not processing speed but memory. The whole model has to fit in GPU memory (VRAM) to run at a usable speed, together with the KV cache that grows with context length and the number of concurrent requests. Quantisation reduces the memory per parameter, and 4-bit quantisation is the common starting point for self-hosted deployments because the quality loss is small for most business tasks.

Model sizeApproximate weights at 4-bitTypical hardwareTypical use
7-8B parametersAbout 5 GBA single GPU with 12-16 GB VRAMClassification, extraction, simple internal Q&A
12-14B parametersAbout 8-10 GBA single GPU with 16-24 GB VRAMRAG over internal documents, summarisation
27-32B parametersAbout 18-20 GBA 24 GB GPU for light use, 48 GB for headroomHigher-quality answers, longer documents
70B parametersAbout 40 GB or moreTwo GPUs or one 80 GB data-centre GPUDemanding reasoning where an API is not allowed

These figures cover the weights only. Leave headroom for the KV cache, especially for retrieval-augmented generation, where each request carries several retrieved passages and the context window fills quickly. A model that fits exactly with a short prompt will slow down or fail once real documents are attached.

A quick check: send the longest realistic prompt several times at once, matching the number of users you expect. If VRAM holds and response times are acceptable, the GPU is sized correctly.

Ollama or vLLM: Choosing the Serving Layer

Ollama packages model download, quantised weights, and an HTTP API into a single binary. It is the fastest way to get a model answering requests, works well for development, internal tools, and moderate traffic, and exposes an OpenAI-compatible endpoint so application code can switch between a local model and an API with minimal change.

vLLM is designed for throughput. Continuous batching and PagedAttention let it serve many concurrent requests from the same GPU far more efficiently, which matters once a chatbot or document pipeline is used by a whole organisation. It also offers an OpenAI-compatible server. A common path is to prototype on Ollama and move the production endpoint to vLLM when concurrency grows, keeping the application code unchanged because both speak the same API shape.

  • Choose Ollama when the priority is a quick, low-maintenance setup for a team or a single application with modest concurrency.
  • Choose vLLM when many users or batch jobs hit the model at the same time and GPU utilisation determines the cost.
  • Keep the application behind an OpenAI-compatible client in both cases, so the serving layer and even the model provider can be swapped later.
  • Pin the model name and quantisation in configuration rather than pulling the latest tag, so a model update is a deliberate change that can be tested and rolled back.

Configure Context Length Deliberately

Ollama starts models with a modest default context length to save memory. That default is fine for chat but too small for retrieval-augmented generation, where the prompt includes instructions, conversation history, and several document passages. When the limit is exceeded, the oldest part of the prompt is silently dropped, which often means the instructions disappear and answer quality falls without any error. Set the context length explicitly for each model, through the num_ctx parameter or a Modelfile, and confirm the memory impact on the target hardware.

Securing a Self-Hosted LLM API

Ollama listens on localhost by default and its API has no authentication. That is safe on a developer laptop and unsafe the moment the bind address is changed to make it reachable from other machines. Anyone who can reach the port can run prompts, pull models, and consume GPU capacity.

  • Keep the model server on a private network and put it behind a reverse proxy or API gateway that enforces authentication and TLS.
  • Expose the model only to the application services that need it, not to every machine on the network.
  • Apply rate limits per client so one runaway batch job cannot starve interactive users of GPU capacity.
  • Log request metadata such as caller, model, token counts, and latency, and decide explicitly whether prompt content is logged, because prompts may contain the very data the private deployment exists to protect.
  • Treat model files as dependencies: download them from trusted sources, record versions, and include them in the same change process as application code.

Evaluate on Your Own Data Before Committing

Public benchmarks say little about how a model handles your documents, your terminology, and Indonesian or mixed-language input. Before choosing a model, collect fifty to a hundred real questions or documents with the answers a domain expert considers correct. Run each candidate model against that set, score the results, and record latency on the target hardware.

The same evaluation set becomes a regression test. When the model, quantisation, prompt, or retrieval settings change, run the set again and compare. Without it, every upgrade is a guess, and quality problems are discovered by users instead of by the team.

How to Roll Out a Private LLM Step by Step

  1. 1Define the use case and the data it touches, then decide whether that data may be sent to an external API. This decision determines whether self-hosting is required or simply one option.
  2. 2Build an evaluation set from real questions and documents, with expected answers reviewed by someone who knows the domain.
  3. 3Test two or three open-weight models of different sizes on Ollama, measure quality and latency, and pick the smallest model that meets the bar.
  4. 4Size the GPU for the chosen model, the required context length, and the expected concurrency, with headroom for growth.
  5. 5Deploy the serving layer on a private network behind authentication, with explicit context length, pinned model versions, and rate limits.
  6. 6Add monitoring for GPU memory, request latency, error rate, and token throughput, and rerun the evaluation set before every model or prompt change.

After Go-Live

New open-weight models are released almost every month, and the one you deploy today will probably be replaced within a year. The parts that last are the evaluation set, the retrieval layer, the access controls, and the monitoring. Once those are in place, switching models means running the evaluation set, comparing the results, and releasing the change like any other.

Key takeaways

  • Self-host when data cannot leave the organisation or when steady volume makes per-token pricing expensive. Otherwise, a managed API is often the faster path.
  • Size GPU memory for the model weights plus the KV cache at realistic context length and concurrency, not for a single short prompt.
  • Start on Ollama, move to vLLM when concurrency grows, and keep application code on an OpenAI-compatible client so the switch is cheap.
  • Never expose the Ollama API directly. Put it on a private network behind authentication, rate limits, and logging.
  • Choose and upgrade models with an evaluation set built from your own questions and documents.

Related articles

More articles on software development, AI, cloud, and infrastructure.

A smartphone showing a folder of messaging apps
AI Development|

AI Chatbots for Business: Use Cases, Channels, and Costs

A guide for business owners considering an AI chatbot: how it differs from a rule-based bot, when it pays off, examples by industry, choosing between website, Telegram, and WhatsApp, what drives the cost, and how to measure the results.

A Linux terminal showing an Ubuntu prompt with the sudo command
Developer Tools|

Setting Up WSL 2 for Software Development on Windows

A practical guide to developing on Windows with WSL 2: installation, where to keep project files, VS Code and Git, limiting memory with .wslconfig, Docker and systemd, networking, backups, and fixes for common problems.

Looking for a software development partner?

Tell us about your project, what you need to build, and the challenges you are facing. We can discuss the technical approach, scope, timeline, and estimated cost.

Start a conversation