A private LLM is a language model that runs on infrastructure the company controls, so prompts, documents, and answers never leave its network. Open-weight models such as Gemma, Qwen, Llama, and DeepSeek have made this practical, and Ollama has made it simple to start. Getting a demo running takes an afternoon. Making it something the business uses every day takes longer, and this guide covers the decisions that come up along the way.
When a Private LLM Makes Sense
Self-hosting is not automatically the better option. Managed APIs such as Claude, GPT, and Gemini offer stronger reasoning, no hardware to buy, and a fast path to a proof of concept. A private LLM earns its place when the data cannot leave the organisation, when request volume is high enough that per-token pricing becomes the dominant cost, or when the system must keep working without an internet connection.
| Factor | Self-hosted LLM | Managed API |
|---|---|---|
| Data location | Stays on servers you control | Sent to the provider, usually under zero-retention terms |
| Reasoning quality | Good for focused tasks such as extraction, classification, and grounded Q&A | Strongest available for complex, open-ended reasoning |
| Cost model | Upfront GPU investment or rental, flat cost per month | Pay per token, grows with usage |
| Operations | Your team runs, patches, and monitors the service | Provider handles capacity and uptime |
| Best fit | Regulated data, internal documents, steady high volume | Prototypes, low volume, tasks needing frontier-grade answers |
Many production systems end up hybrid. Sensitive documents are processed by a self-hosted model, while non-sensitive tasks that benefit from stronger reasoning are routed to an API. Designing that routing early is easier than retrofitting it after the first model has been integrated everywhere.
Sizing GPU Memory for an Open-Weight Model
The first hardware question is not processing speed but memory. The whole model has to fit in GPU memory (VRAM) to run at a usable speed, together with the KV cache that grows with context length and the number of concurrent requests. Quantisation reduces the memory per parameter, and 4-bit quantisation is the common starting point for self-hosted deployments because the quality loss is small for most business tasks.
| Model size | Approximate weights at 4-bit | Typical hardware | Typical use |
|---|---|---|---|
| 7-8B parameters | About 5 GB | A single GPU with 12-16 GB VRAM | Classification, extraction, simple internal Q&A |
| 12-14B parameters | About 8-10 GB | A single GPU with 16-24 GB VRAM | RAG over internal documents, summarisation |
| 27-32B parameters | About 18-20 GB | A 24 GB GPU for light use, 48 GB for headroom | Higher-quality answers, longer documents |
| 70B parameters | About 40 GB or more | Two GPUs or one 80 GB data-centre GPU | Demanding reasoning where an API is not allowed |
These figures cover the weights only. Leave headroom for the KV cache, especially for retrieval-augmented generation, where each request carries several retrieved passages and the context window fills quickly. A model that fits exactly with a short prompt will slow down or fail once real documents are attached.
A quick check: send the longest realistic prompt several times at once, matching the number of users you expect. If VRAM holds and response times are acceptable, the GPU is sized correctly.
Ollama or vLLM: Choosing the Serving Layer
Ollama packages model download, quantised weights, and an HTTP API into a single binary. It is the fastest way to get a model answering requests, works well for development, internal tools, and moderate traffic, and exposes an OpenAI-compatible endpoint so application code can switch between a local model and an API with minimal change.
vLLM is designed for throughput. Continuous batching and PagedAttention let it serve many concurrent requests from the same GPU far more efficiently, which matters once a chatbot or document pipeline is used by a whole organisation. It also offers an OpenAI-compatible server. A common path is to prototype on Ollama and move the production endpoint to vLLM when concurrency grows, keeping the application code unchanged because both speak the same API shape.
- Choose Ollama when the priority is a quick, low-maintenance setup for a team or a single application with modest concurrency.
- Choose vLLM when many users or batch jobs hit the model at the same time and GPU utilisation determines the cost.
- Keep the application behind an OpenAI-compatible client in both cases, so the serving layer and even the model provider can be swapped later.
- Pin the model name and quantisation in configuration rather than pulling the latest tag, so a model update is a deliberate change that can be tested and rolled back.
Configure Context Length Deliberately
Ollama starts models with a modest default context length to save memory. That default is fine for chat but too small for retrieval-augmented generation, where the prompt includes instructions, conversation history, and several document passages. When the limit is exceeded, the oldest part of the prompt is silently dropped, which often means the instructions disappear and answer quality falls without any error. Set the context length explicitly for each model, through the num_ctx parameter or a Modelfile, and confirm the memory impact on the target hardware.
Securing a Self-Hosted LLM API
Ollama listens on localhost by default and its API has no authentication. That is safe on a developer laptop and unsafe the moment the bind address is changed to make it reachable from other machines. Anyone who can reach the port can run prompts, pull models, and consume GPU capacity.
- Keep the model server on a private network and put it behind a reverse proxy or API gateway that enforces authentication and TLS.
- Expose the model only to the application services that need it, not to every machine on the network.
- Apply rate limits per client so one runaway batch job cannot starve interactive users of GPU capacity.
- Log request metadata such as caller, model, token counts, and latency, and decide explicitly whether prompt content is logged, because prompts may contain the very data the private deployment exists to protect.
- Treat model files as dependencies: download them from trusted sources, record versions, and include them in the same change process as application code.
Evaluate on Your Own Data Before Committing
Public benchmarks say little about how a model handles your documents, your terminology, and Indonesian or mixed-language input. Before choosing a model, collect fifty to a hundred real questions or documents with the answers a domain expert considers correct. Run each candidate model against that set, score the results, and record latency on the target hardware.
The same evaluation set becomes a regression test. When the model, quantisation, prompt, or retrieval settings change, run the set again and compare. Without it, every upgrade is a guess, and quality problems are discovered by users instead of by the team.
How to Roll Out a Private LLM Step by Step
- 1Define the use case and the data it touches, then decide whether that data may be sent to an external API. This decision determines whether self-hosting is required or simply one option.
- 2Build an evaluation set from real questions and documents, with expected answers reviewed by someone who knows the domain.
- 3Test two or three open-weight models of different sizes on Ollama, measure quality and latency, and pick the smallest model that meets the bar.
- 4Size the GPU for the chosen model, the required context length, and the expected concurrency, with headroom for growth.
- 5Deploy the serving layer on a private network behind authentication, with explicit context length, pinned model versions, and rate limits.
- 6Add monitoring for GPU memory, request latency, error rate, and token throughput, and rerun the evaluation set before every model or prompt change.
After Go-Live
New open-weight models are released almost every month, and the one you deploy today will probably be replaced within a year. The parts that last are the evaluation set, the retrieval layer, the access controls, and the monitoring. Once those are in place, switching models means running the evaluation set, comparing the results, and releasing the change like any other.
Key takeaways
- Self-host when data cannot leave the organisation or when steady volume makes per-token pricing expensive. Otherwise, a managed API is often the faster path.
- Size GPU memory for the model weights plus the KV cache at realistic context length and concurrency, not for a single short prompt.
- Start on Ollama, move to vLLM when concurrency grows, and keep application code on an OpenAI-compatible client so the switch is cheap.
- Never expose the Ollama API directly. Put it on a private network behind authentication, rate limits, and logging.
- Choose and upgrade models with an evaluation set built from your own questions and documents.


