Introduction
A self-hosted LLM is a language model you run on infrastructure you control, instead of renting access to someone else's model through a per-token API. For most companies in 2026, that is still the wrong default: hosted APIs are cheaper to start, improve every few months, and need no GPU expertise. But there is a specific kind of workload where owning the model wins, and I think more teams have one than realise it.
This guide is for CTOs and founders weighing whether to stop paying per token. I will cover when self-hosting LLMs pays off and when it does not, the hardware and tools involved, from Ollama on a laptop to production serving, and the one case where we did it ourselves: a fine-tuned open-weight model that does a single narrow task, passes every case in its evaluation gate, and runs behind our own endpoint.
The short version of our view: self-hosting pays off when a task is narrow and repeated at volume, when the output format is fixed, or when data cannot leave your own infrastructure. It does not pay off for open-ended chat or low volume.
What a Self-Hosted LLM Means in 2026
The term covers more ground than it used to, so it is worth being precise. A self-hosted LLM always runs on weights you have downloaded and control, but where those weights run varies, and so do the trade-offs in cost, privacy and effort.
Open-Weight Is Not the Same as Open Source
Most LLMs people self-host are open weight models: the trained weights are published under a licence that lets you run and modify them, but the training data and full training code usually are not. Google's Gemma 4, Meta's Llama, Mistral's models and OpenAI's gpt-oss are in this category. Calling them open source LLMs is common but loose; the licence is what matters for a business.
Licensing and Commercial Use
Licences differ more than the marketing suggests. Google's Gemma 4 model card lists Apache 2.0, which permits commercial use and redistribution of fine-tuned derivatives. Llama models ship under Meta's own community licence with its own conditions. Read the licence before you build on a model, especially if you plan to fine-tune it and offer the result to clients.
Three Ways to Host Your Own Model
There are three common setups for hosting LLMs, from most to least control:
- On-premise AI on your own servers in a premise data center or private cloud, where nothing leaves your network. Maximum control, maximum operational burden.
- A managed GPU host that runs your weights behind an endpoint you own. You control the model and its configuration; the host handles hardware and scaling.
- Open weights on a hyperscaler's model service, such as Amazon Bedrock, which now hosts several open-weight models. You get open weights with cloud-provider billing, but you choose from their catalogue rather than serving your own fine-tune.
We use the second setup for our own model and have used the third in client work. On a document extraction project for a fintech client, we tested gpt-oss-20b on Bedrock alongside hosted proprietary models; it was cheap per document but produced one invalid output in 85, so it was not chosen for production.
Local LLMs Are a Different Thing
Searches for a local LLM and how to run an LLM locally are mostly about running a model on your own machine, for privacy, offline use or experimentation. That is a good way to learn how models behave, and it is fine for personal tools. It is not a production architecture: a business deployment needs an always-on endpoint, monitoring, a versioned serving configuration and an evaluation gate, which is what the rest of this guide covers.
Is Self-Hosting an LLM Worth It? When It Wins and When APIs Win
The decision is rarely about model quality alone. It is about the shape of the workload, and five factors decide it.
| Factor | Self-hosting tends to win | A hosted API tends to win |
|---|---|---|
| Task width | One narrow, repeated task with a fixed output format | Open-ended chat, research, varied requests |
| Volume | High, steady volume that keeps a GPU busy | Low or spiky volume |
| Data privacy | Sensitive data that legally or contractually cannot leave your infrastructure | Data you may send to a processor under a standard agreement |
| Model change | You want the model frozen until you choose to change it | You want the newest model as soon as it ships |
| Team skills | You can own serving, quantisation and evals | You would rather pay someone else to |
The first row matters most. A model that has to handle anything needs to be as capable as possible, and the best general models are proprietary. A model that has to do one thing well can be much smaller, and a smaller model on your own endpoint is where the economics and control start to favour self-hosting. Our comparison of small language models vs large ones shows how far a small model can go on a constrained task.
Volume is the second filter. A GPU endpoint costs money whether requests arrive or not, so self-hosting only pays when traffic keeps it busy, or when the other factors (privacy, control) justify the idle time. Cost control with a self-hosted model means managing utilisation, not counting tokens.
Hardware Requirements: How Much VRAM a Self-Hosted LLM Needs
Hardware is the first practical question, and it comes down to GPU memory. The model's weights must fit in VRAM, with room left for the context the model is processing.
The Memory Rule of Thumb
Each parameter takes about two bytes at 16-bit precision (BF16), one byte at 8-bit (INT8) and about half a byte at 4-bit. So the weights alone need roughly:
| Model | Parameters | 16-bit | 8-bit | 4-bit |
|---|---|---|---|---|
| Mistral 7B | 7B | about 14 GB | about 7 GB | about 4 GB |
| Llama 3.1 8B | 8B | about 16 GB | about 8 GB | about 4 GB |
| Gemma 4 26B-A4B | 25.2B | about 50 GB | about 25 GB | about 13 GB |
| Llama 3.1 70B | 70B | about 140 GB | about 70 GB | about 35 GB |
These are approximations for the weights only. Mixture-of-experts models such as Gemma 4 26B-A4B still need all their parameters in memory, even though only 3.8 billion are active per token, so they are fast but not small in VRAM.
Context Adds to the Requirements
The context window needs its own memory, the key-value cache, which grows with the length of the input and the number of requests served at once. A model that fits comfortably for short prompts can run out of VRAM on long documents or under concurrent load. Size your hardware for the longest realistic request at peak concurrency, not for a single short test.
Quantisation Trades Memory for Precision
Quantisation stores weights at lower precision to cut memory. 8-bit usually costs little quality; 4-bit saves more memory with a larger risk, which you should measure on your own test cases rather than assume. We publish our own model in BF16 and serve it at INT8, which halves memory against BF16.
Small LLMs can run on a CPU or a laptop with enough RAM, especially at 4-bit, but they generate text far more slowly than on a GPU. That is fine for experiments on your own machine and too slow for most production workloads.
Tools for Self-Hosting LLMs: From a Laptop to Production
The tools for self-hosting split into two groups: tools for running LLMs on a single machine, and serving engines built for production traffic.
Ollama: The Fastest Local Setup
Ollama is the most common way to run an LLM locally. It downloads LLMs, handles quantised formats and exposes a local API on port 11434, including an OpenAI-compatible endpoint. Our own open-source MDMA CLI ships with an Ollama preset pointing at http://localhost:11434/v1, so developers can try MDMA against a local model with no API key.
The official Docker image is the quickest setup on a machine with an NVIDIA GPU, as described on the Ollama Docker Hub page:
docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
docker exec -it ollama ollama run llama3Drop --gpus=all to run it on CPU only. The named volume keeps downloaded models between container restarts, so you only pull each model once.
Open WebUI: A Web Interface for Your Models
Open WebUI adds a ChatGPT-style web interface on top of Ollama or any OpenAI-compatible endpoint, which is useful for letting non-developers try a model. Open WebUI's quick start runs it as a Docker container that connects to Ollama on the host:
docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway \
-v open-webui:/app/backend/data --name open-webui --restart always \
ghcr.io/open-webui/open-webui:mainThe web interface then opens at http://localhost:3000. For a team pilot, add authentication and a secret key before anyone outside your machine can reach it.
LM Studio and LocalAI
LM Studio is a desktop application for downloading and running LLMs with a graphical interface, and it can serve them through an OpenAI-compatible local server on port 1234. Our MDMA demo lists LM Studio alongside Ollama as a supported local provider. LocalAI is an MIT-licensed, OpenAI-compatible runtime that puts CPU support first and runs as a single container, which makes it a reasonable choice when you have no GPU at all.
vLLM for Production Serving
Local tools are built for one user. For production, you need a serving engine that batches many concurrent requests efficiently. vLLM is the most widely used open source option: vLLM's quickstart starts an OpenAI-compatible server on port 8000 with a single vllm serve command, and it ships an official Docker image. We serve our own fine-tuned model with vLLM on a managed GPU platform, and its model card recommends vLLM or SGLang for OpenAI-compatible inference.
Kubernetes and Managed Platforms
At larger scale, teams run serving engines as containers on Kubernetes with GPU nodes, which gives autoscaling and rolling upgrades at the cost of significant operational work. A managed GPU platform is the lighter alternative: you bring the model and its configuration, and the platform runs the containers and GPUs. That is the route we use for our own model.
Our Case: One Narrow Task, One Fine-Tuned Model
We self-host exactly one model, for exactly one job. Our open-source MDMA format lets a model return interactive components, such as forms, tables and approval steps, inside ordinary Markdown. The job is to turn a compact description of an interface into a valid MDMA document, every time.
Why This Task Fits
That description is written in MDMA-IL, a one-line-per-component intent language. A line like form#contactfull-name:t, email^:e means a contact form with a required name field and a required, sensitive email field, submitted to a named action. The model's output must be valid YAML inside fenced blocks that our validator accepts, with no room for creative interpretation.
That is the ideal self-hosting workload: one input grammar, one output format, a deterministic validator that can check every answer, and volume that grows with every product that embeds MDMA. A general frontier model can do it, but it is paying for capabilities the task never uses.
What We Built
We fine-tuned Google's Gemma 4 26B-A4B, a mixture-of-experts model with 25.2 billion total parameters of which only 3.8 billion are active for each token, according to Google's model card. The fine-tune is a LoRA adapter trained with Unsloth and merged back into the base weights, and we published the result as MobileReality/mdma-gemma4-26b-dsl-unsloth-v1 under Apache 2.0, as we described when we released the open-source generative UI model. The training ran on a single NVIDIA A100 80GB GPU.
We chose a mixture-of-experts base because it gives a large model's knowledge at a small model's per-token compute. Our full account of the training decisions is in our write-up on fine-tuning an open-weight model for generative UI.
LLM Deployment in Production: What It Takes to Serve the Model
Training gets the attention, but serving is where self-hosting succeeds or fails. A fine-tuned model that is unstable under real settings is worse than an API, because nobody else is on call for it.
Quantise for Serving, Publish at Full Precision
We publish the merged model in BF16, the precision it was trained in, so anyone can re-quantise it for their own hardware. We serve it INT8-quantised with vLLM behind an OpenAI-compatible endpoint on Modal, a serverless GPU platform. INT8 roughly halves the VRAM needed compared with BF16, which decides how much GPU you rent.
The OpenAI-compatible endpoint matters more than it looks. Every client, test harness and SDK that speaks the OpenAI chat format works against it unchanged, so swapping between our model and a hosted one is a configuration change, not a rewrite. The same property is why Ollama, LM Studio, LocalAI and vLLM all expose OpenAI-style APIs.
Write Down the Serving Contract
The most useful artefact of the whole project is a short list of serving settings, because Gemma 4 has a failure mode that can look like a broken fine-tune. With reasoning enabled, it can fall into repetitive thinking loops and never produce the answer. Our evaluated configuration turns thinking off with enablethinking: false and adds a repetition guard, minp: 0.02 and repetition_penalty: 1.1, with output capped at 2,048 tokens.
Treat those settings as part of the model, not as client options. We learned that the hard way: as the model moved from an earlier, smaller version to the current one, comments in our own configuration files and the recommended temperature drifted out of sync with what the evaluation ran. The fix is to keep one serving contract, versioned next to the model, and to test against exactly that contract.
Serving on Modal: What Cold Starts Taught Us
Our model runs on Modal, a serverless GPU platform. Serverless sounds like the perfect fit for a self-hosted LLM: you pay only while a container runs, and it scales to zero when nobody is using it. The catch is the cold start, and for a large model it is the single biggest operational trade-off we have run into.
Why We Chose Modal
The reason was practical rather than strategic. Modal's pricing bills GPU time per second and includes 30 dollars of free compute credits a month on its Starter plan, which let us fine-tune, serve and test the model without a separate budget line. For an experiment that might not reach production, paying nothing while the endpoint is idle mattered more than squeezing out the last second of latency.
How Long a Cold Start Takes
A cold start happens when a request arrives and no container is running, so the platform has to start one and load the model before it can answer. For our models, that is not a matter of seconds:
| Model served | Weights loaded at start | Typical cold start |
|---|---|---|
| Gemma 4 26B-A4B (our fine-tune, INT8) | about 52 GB | 3 to 4 minutes, sometimes 6 to 8 |
| Gemma 4 31B dense, before optimisation | about 62 GB | 6 to 10 minutes |
| Gemma 4 31B dense, with AWQ quantisation and Marlin kernels | smaller quantised weights | about 3 minutes to the first successful response |
The spread on the 26B model is real: the same container image sometimes came up in under four minutes and sometimes took twice as long. If a user-facing feature depends on that endpoint, a multi-minute first response is not a slow request, it is an outage.
What Makes Cold Starts Slow
Two things dominate. First, every new container reads 52 to 62 GB of weights from a storage volume and loads them into GPU memory. Second, vLLM compiles CUDA graphs for faster inference when eager mode is off, which adds its own cost measured in minutes. We keep vLLM's compilation cache (VLLMCACHEROOT) on a persistent volume, so only the first cold start pays for compilation; it helps, but it does not make startup fast.
Quantisation shortens the load by shrinking the weights, which is part of why the AWQ version of the 31B model started much faster than the unoptimised one.
Warm Containers Cost Money, Cold Ones Cost Time
Modal's cold start guide offers the standard levers. mincontainers keeps a floor of running containers so the function never scales to zero, and scaledownwindow keeps an idle container alive for longer before shutting it down (60 seconds by default, up to 20 minutes). Modal also offers memory snapshots to reduce startup work.
In practice the trade-off is blunt. Either you keep a container running and pay for the GPU all the time, or you tune scaledown_window so it stays alive longer, which again means paying for more GPU time. Serverless GPU hosting does not remove the cost of keeping a large model ready; it only lets you choose who pays for it, you in GPU hours or your users in waiting time.
For experiments, internal tools and batch jobs that can tolerate a slow first request, scale-to-zero hosting is excellent value. For a customer-facing feature with a latency expectation, budget for at least one warm container, or use a provider that keeps models loaded for you. Put that cost into the comparison with per-token API pricing before you decide to self-host.
Alternatives to Modal: Parasail and Similar Platforms
Modal is not the only option between running your own GPUs and paying per token. We first heard about Parasail from a company evaluating MDMA for its own model routing, which serves models on it. According to Parasail's documentation, it offers serverless endpoints for open-weight models behind an OpenAI-compatible API, plus private dedicated GPU endpoints for your own model with full control over model, hardware and scaling, billed per GPU-hour.
The difference matters for cold starts. A shared serverless endpoint for a popular open model is usually already warm because other customers keep it busy, while a dedicated endpoint for your own fine-tune behaves more like a warm container you pay for by the hour. We have not run our model on Parasail, so treat this as a description of the options, not a recommendation from experience.
Proving It Works Before You Switch: An Eval Gate, Not a Demo
A self-hosted model should replace an API only after it passes the same tests the API would have to pass. A demo proves the model can work; a gate proves it works on cases it has never seen.
Our Gate and Its Results
Our gate holds 95 scenarios that were kept out of training. Every output must pass our MDMA validator, with no partial credit. The current 26B model passes all 95, and it also passes the author, custom-prompt, fixer, flow and tool-guidance suites we run against it: 181 of 181 committed cases in total. On the holdout run, the median generation took about 1.4 seconds end to end.
The earlier attempt is the useful comparison. A fine-tune on the much smaller Gemma 4 E4B, served with a 2,048-token context, topped out at around 90.5% valid on the same holdout. Without a gate, that model would have looked fine in a demo and failed about one request in ten in production. Our article on how we run LLM evaluation before shipping covers the suites in detail.
The gate also tells you when not to switch. If your self-hosted candidate scores below the hosted model you use today on your own cases, the privacy and control arguments rarely make up for it.
Private LLM, Data Privacy and Security
For many companies the strongest argument for a private LLM is not cost but control over where sensitive data goes. When a model runs on your own infrastructure, prompts and outputs never reach a third-party model provider or external servers, which removes a processor from your data map and a dependency from your risk register.
This matters most in regulated sectors. Fintech and proptech clients regularly restrict sending customer personal data to external API providers, and a self-hosted model behind a private endpoint keeps interactive AI features available without that transfer. In the EU, it also fits the broader push for control over infrastructure and data that we cover in our article on digital sovereignty for European companies.
Be precise about what self-hosting does and does not solve. It keeps data away from model providers, but your GPU host is still a processor, so the hosting contract, region and access controls still need a security review. On-premise hosting removes that last processor, at the cost of running the hardware yourself.
Zero-data-retention options and regional endpoints cover many privacy requirements without self-hosting, and the shared inference library we use in client work requests zero data retention on every OpenRouter fallback call. Self-hosting is the answer when contracts or regulators require that the data never leaves, not merely that it is not stored.
Quick Answers on Self-Hosting LLMs
These are the questions we hear most often from teams considering a self-hosted LLM.
How Much Does It Cost to Self-Host an LLM?
The cost is driven by GPU memory and utilisation, not by tokens. A model that fits on one GPU at 8-bit precision is far cheaper to host than one that needs several, and an endpoint that sits idle costs the same as one that is busy.
Add the engineering time for serving, monitoring and evaluation, which often exceeds the hardware bill for a single model. Our own experiments have so far fitted inside the free monthly credits of a serverless GPU platform, but a production endpoint that has to stay warm is a different budget. We do not yet publish a per-request cost for our own model, so measure your current API cost per task before you compare.
Which LLM Model Is Best for Self-Hosting?
There is no single best model, only the best model for your task and hardware. Start with the smallest open-weight model that passes your own evaluation cases, with a licence that allows your commercial use. For narrow, structured tasks, a fine-tuned small or mixture-of-experts model often beats a larger general one; for broad tasks, you may need a 70B-class model and the hardware that goes with it.
What LLM Can I Run at Home?
On a laptop or a desktop with a consumer GPU, 7B to 8B models such as Mistral 7B or Llama 3.1 8B run comfortably at 4-bit through Ollama or LM Studio. Larger LLMs need more VRAM than most consumer cards provide. That is enough to learn how models behave and to prototype, but plan for proper serving hardware before production.
Is Self-Hosting an LLM Worth It?
It is worth it when you have a narrow, high-volume task, a fixed output format, or sensitive data that cannot leave your infrastructure, and a team that can own serving and evaluation. For everything else, a hosted API is usually cheaper, better and less work. The sections above give you the test.
When Not to Self-Host, and the Hybrid Approach
Self-hosting is a commitment, and most teams should not make it yet. Skip it if any of these apply:
- Your task is open-ended. A general assistant, a research agent or a support bot that must handle anything needs the best general model, and that is usually a hosted one.
- Your volume is low or unpredictable. An idle GPU endpoint costs more than the API calls it replaces.
- You have no evaluation set. Without your own test cases, you cannot prove the self-hosted model is good enough, and you will not notice when it degrades.
- Nobody owns serving. Quantisation, endpoint configuration, cold starts and upgrades are ongoing work, not a one-off task.
- You are chasing savings you have not measured. Measure your current cost per task first, then compare.
Most teams get the best of both with a hybrid approach: route narrow, high-volume steps to a small or self-hosted model, and keep a strong hosted model for everything else. Because every option here speaks the same OpenAI-compatible API, moving a single step between them is a configuration change, which keeps the decision reversible.
Matt Sadowski
CEO of Mobile Reality
Own the Model Behind Your AI Feature
We help teams decide whether a smaller, fine-tuned or self-hosted model can take over a narrow, high-volume task, and prove it before anything changes in production.
Fine-tuning open-weight models for narrow tasks with a fixed output format.
A held-out evaluation gate on your own cases before any switch.
Serving behind an OpenAI-compatible endpoint, quantised for your hardware.
Data that stays on infrastructure you control.
An honest answer when a hosted API is still the better choice.
Conclusion
A self-hosted LLM is the right call for a specific workload, not a general upgrade over hosted APIs. Our own experience with one narrow, fine-tuned model points to a simple test: own the model when the task is narrow, the format is fixed, the volume is steady or the data cannot leave, and keep paying per token otherwise.
- Open-weight models under permissive licences, such as Gemma 4 under Apache 2.0, make self-hosting a legal and practical option for commercial use.
- Hardware requirements come down to VRAM: about two bytes per parameter at 16-bit, one at 8-bit, plus room for context.
- Ollama, LM Studio and Open WebUI are the fastest way to run an LLM locally; production needs a serving engine such as vLLM or a managed GPU platform.
- Switch only after the model passes a held-out evaluation gate on your own cases, and version the serving contract with the model.
- Data privacy is often the real reason to self-host, but zero-data-retention APIs cover many requirements without it.
If you have a narrow, repeated task that you suspect a smaller owned model could handle, we can help you test that on your own cases before you commit, through our custom AI model development services.
Frequently Asked Questions
When does a self-hosted LLM make sense for a business?
When the task is narrow and repeated at volume, the output format is fixed, or the data cannot leave your infrastructure. For open-ended chat, research or low and unpredictable volume, a hosted API is usually cheaper and better. Many teams never reach the point where self-hosting pays off, and that is fine.
What is the difference between an open-weight and an open-source model?
An open-weight model publishes its trained weights under a licence that lets you run and modify them, but usually not the training data or full training code. Gemma 4, for example, is released under Apache 2.0, which allows commercial use and fine-tuned derivatives. For business use, the licence terms matter more than the label.
Is a self-hosted LLM cheaper than an API?
Not automatically. A GPU endpoint costs money whether requests arrive or not, so it only pays off when traffic keeps it busy or when control and data residency justify the idle time. Measure your current cost per task before assuming a saving; we do not publish a per-request comparison for our own model yet.
How do you know a self-hosted model is good enough to replace an API?
Run it through an evaluation gate of cases it never saw in training, with automatic checks on every output. Our fine-tuned Gemma 4 model passes all 95 held-out cases in its gate, while an earlier, smaller version topped out around 90.5%, which would have meant roughly one failed request in ten.
More on AI Cost, Small Models and Evaluation
What an AI system costs, which model it runs on, and how you prove it still works are one decision, not three. These articles cover the build and run cost, model choice and the evaluation that makes switching safe:
- LLM Router: How We Route Models by Cost, Quality and Failure
- Prompt Caching: A Practical Guide for Claude, OpenAI, Gemini
- SLM vs LLM: Key Differences and When Small Models Win
- Fine-Tuning LLMs: A Guide From Our Gemma 4 Case Study
- AI Development Cost in 2026: What Drives It and Real Ranges
- LLM Evaluation in Practice: How We Test Across 32 Models
- Context Is King: How to Serve It With Context Engineering
- Structured LLM Output Without JSON Schemas | MDMA
- Build AI Agents with 75+ Deployments Cutting Costs 60% in 2026
Want to know whether a smaller or self-hosted model could handle part of your workload? Our custom AI model development team starts with an evaluation on your own cases.
