Introduction
Fine-tuning an LLM means continuing the training of an existing model on your own examples, so it learns a behaviour that prompting alone does not produce reliably.
In 2026 it is easier and cheaper than ever, and still the wrong tool for most problems. Prompting and retrieval solve the majority of business use cases faster, and they keep you on the newest models.
This article is for engineering leads deciding between prompting, retrieval-augmented generation (RAG) and tuning the model. I will explain which problem each one solves, then walk through the one fine-tune we run in production: a LoRA fine-tune of Gemma 4 that turns a compact interface description into a valid interactive document, and reaches 95 of 95 on its held-out gate once the instructions and decoding settings match the ones we validated.
Our rule of thumb is simple. Fine-tune when the task is narrow and the output grammar is fixed. In our case, prompting got us most of the way, and the fine-tune plus a strict evaluation gate closed the rest.
Prompting vs RAG vs Fine-Tuning: Which Problem Each One Solves
These three techniques are often presented as rungs on a ladder. They are not. Each one fixes a different kind of gap, and picking the wrong one wastes weeks.
Prompting Fixes Instructions
If the model can already do the task but does not know what you want, give it better instructions. Clear instructions, a defined output format and a few examples solve most problems, and they cost nothing to change.
The debate about fine tuning vs prompt engineering usually ends here: if a frontier model passes your tests with good instructions, you do not need a fine-tune.
RAG Fixes Knowledge
If the model lacks facts, such as your product documentation, contracts or yesterday's prices, give it those facts at request time. RAG retrieves the relevant documents and puts them in the request.
On RAG vs fine tuning, the rule is that knowledge which changes belongs in retrieval, because a fine-tuned model only knows what it saw during training and has to be retrained to learn anything new.
Fine-Tuning Fixes Behaviour
Tuning changes how the model responds: its format, style, vocabulary or the mapping from a specific input to a specific output. It is the right tool when you need the same behaviour thousands of times, when that behaviour is hard to describe in words, or when you want a smaller model to do reliably what only a large model does with prompting.
Types of Tuning: SFT, Instruction and Efficient Methods
Tuning comes in several flavours, and the terms overlap. Supervised fine-tuning (SFT) trains on pairs of an input and the output you want, and it is what we did. Instruction tuning is supervised training on instruction and answer pairs, which is how chat models learn to follow requests. Efficient fine-tuning methods, such as LoRA, train a small number of added parameters instead of every weight.
All of them rely on transfer learning: the model keeps what it learned in pretraining and adapts it to your task instead of starting from scratch.
Reinforcement learning from human feedback (RLHF) goes one level further than supervised training. Instead of copying example answers, the model is trained against a reward that reflects human preferences, for example which of two answers people prefer. We did not use reinforcement learning here, because a strict validator already gave us a clear, automatic check of every answer.
Fine-Tuning Is Not Distillation
Model distillation is a related but different technique. In distillation, a smaller "student" model is trained to imitate the outputs of a larger "teacher" model, often across a broad range of tasks, as described in Hinton and colleagues' original paper on knowledge distillation.
What we did is supervised fine-tuning on one narrow task, using a validated dataset of input and output pairs. The goal was task reliability on one input grammar, not breadth across many tasks.
Why We Fine-Tuned: One Input Grammar, One Output Format
Our open-source MDMA format lets a model return interactive components, such as forms, tables, approval steps and charts, inside ordinary Markdown. Getting models to write MDMA correctly is mostly a prompting problem, and we solve it that way for frontier models: our prompt pack includes model-specific prompts, and our test matrix covers 32 model variants from OpenAI, Anthropic, Google and xAI.
The Task: Turning MDMA-DSL Into Valid Documents
For our own model we narrowed the task further. The input is MDMA-DSL, a compact domain-specific language with one line per component.
A line like form#contactfull-name:t, email^:e describes a contact form with a required name field, a required and sensitive email field, and a named submit action. The output must be valid YAML inside fenced Markdown blocks that our validator accepts.
That is about as fixed as a grammar gets. There is one correct structure for each input, a validator that can check every answer without human judgement, and no need for knowledge that changes over time.
Data Preparation: A Small, Validated Dataset
Tuning data is a set of examples, each a pair of an input and the output you want. Ours pair an MDMA-DSL description with a valid MDMA document. We generated 1,560 pairs with a larger teacher model across 52 business domains, and every pair went through our validator at generation time. The quality filter dropped only 2, which left 1,558. We split them into 1,298 training examples and 260 validation examples, about 25 training examples per domain, roughly 84% in English and 16% in Polish.
The dataset is small on purpose. The task is narrow, and every example passed the same validator that checks documents in production. The holdout is separate: 95 hand-curated scenarios, 68 from our catalog and 27 from regression cases, never generated by the teacher model and never used for training.
Why Not Just Keep Prompting?
Prompting a frontier model works for this task, so the fine-tune was not about making it possible. It was about owning it: a model we can run on our own endpoint, keep frozen until we choose to change it, and offer to enterprise teams that cannot send customer data to an external model provider.
A narrow, fixed task is also exactly where a smaller model can match a large one, as our comparison of small language models vs large ones shows.
Choosing the Base Model: Gemma 4 26B-A4B Over E4B
The base model decides the ceiling of a fine-tune. A fine-tune teaches behaviour; it does not add capacity the base model never had.
We Started Small With E4B
We started with Gemma 4 E4B, the small on-device member of the family. It was cheap to train and serve, it learned the grammar, and it reached 94.7% on our lab holdout.
Since a small model already did well, we decided to try a larger and more capable one on the same recipe and see how much further it could go.
Why a Mixture-of-Experts Base
We moved to Gemma 4 26B-A4B, a mixture-of-experts model. According to Google's Gemma 4 model card, it has 25.2 billion total parameters but only 3.8 billion active per token, a 256K-token context window and an Apache 2.0 licence.
That combination was the point: the knowledge of a mid-sized model, the per-token compute of a small one, plenty of context, and a licence that let us publish our fine-tune.
We fine-tuned the instruction-tuned variant, unsloth/gemma-4-26B-A4B-it, which is Unsloth's mirror of the original weights. Starting from an instruction-tuned model means the fine-tune only has to teach the task, not basic instruction following.
A Detour Through the 31B Dense Model
Our first plan was to fine-tune the 26B-A4B with QLoRA. It did not work: the mixture-of-experts layers store their experts as fused 3D tensors, and bitsandbytes cannot quantise those to 4-bit.
So we trained the dense Gemma 4 31B instead. It scored 96.8% on the holdout (92 of 95) when served with AWQ quantisation, at about 29 tokens per second on an A100 40GB.
Quality was fine, but a dense model activates all of its parameters on every token, and that speed is fixed by the model and the hardware, so we could not tune our way out of it.
A slightly lower score is different: it can still be improved on the system prompt side, and MDMA also has a fixer prompt and an auto-fixer that repair documents. So we went back to the 26B-A4B and found a way to train it: 16-bit LoRA with Unsloth, which fits on a single A100 80GB.
LoRA Fine-Tuning With Unsloth, Merged Into the Base
Updating every weight in the model needs a lot of GPU memory and produces a full-size copy for each experiment. LoRA fine tuning avoids both problems.
What LoRA Does
LoRA, short for low-rank adaptation, freezes the original model weights and trains small additional matrices alongside them, as introduced in the LoRA paper from Microsoft researchers.
Only those small matrices are updated, so training needs far less memory and the result is an adapter of a fraction of the model's size. You can train several adapters on the same base and compare them cheaply. Because the original weights stay frozen, LoRA also limits catastrophic forgetting: after training, our model still handled tool calling the way the base model does.
Why Unsloth
We trained with Unsloth, an open-source framework that speeds up training and reduces its memory use for popular open-weight models, Gemma included.
As we stated when we released the model, the fine-tune used 16-bit LoRA on a single NVIDIA A100 80GB GPU. A 26-billion-parameter model fine-tuned on one GPU is a good illustration of how far the cost of tuning has fallen.
Training Hyperparameters
Hyperparameters are the settings you choose before training starts, and for a LoRA run only a few of them matter. We used a rank of 64 with alpha 128 on all linear layers, which meant training about 2 billion trainable parameters, 7.3% of the model, including the large expert layers.
The learning rate was 1e-5 with a cosine schedule and 5% warmup, with weight decay 0.01 and 8-bit AdamW. A batch size of 1 with 16 gradient accumulation steps gave an effective batch of 16, because a larger per-device batch did not fit on the card.
We trained for two epochs with a maximum sequence length of 4,096 tokens. Validation loss had flattened out by the end of the second epoch, so a third one would not have helped.
The run took about two hours and peaked at 74.4 GB of the card's 80 GB, so a single A100 was close to the limit. The GPU time cost roughly 5 to 7 dollars.
LoRA, Not QLoRA
QLoRA is a popular variant that loads the base model in 4-bit precision during training to save even more memory, described in the QLoRA paper. It makes large fine-tunes possible on smaller GPUs, at some risk to quality.
We could not use it here: bitsandbytes cannot quantise the fused expert tensors of this MoE model. Our run used 16-bit LoRA instead, so the base weights stayed at higher precision during training.
Merging the Adapter
After training, we merged the LoRA adapter back into the base model weights, producing a single standalone model. A merged model is simpler to serve: any inference stack that can run Gemma 4 can run it, with no adapter loading.
We published it on Hugging Face as MobileReality/mdma-gemma4-26b-dsl-unsloth-v1 under Apache 2.0. The training data and serving configuration are not part of that release.
The Gate: 95 Held-Out Cases and a Validator
A fine-tune is only as good as the test that decides whether it ships. We built the gate around the same validator that checks every MDMA document in production.
A Holdout Set the Model Never Saw
We kept 95 scenarios out of training entirely. Each one is an MDMA-DSL description, and the model's answer must pass our MDMA validator with no partial credit. Scoring against a held-out set is the only way to know the model learned the grammar rather than memorising its training examples.
Promptfoo Runs the Suites
The gate runs on promptfoo, the same open-source evaluation tool we use for every prompt in the project. Beyond the holdout set, we run the fine-tuned model through suites for document authoring, custom system prompts, repairing broken documents, interactive flows and tool-calling guidance.
Our write-up on how we run LLM evaluation across models before shipping covers the suites and assertions in detail. In short, the tools and frameworks in the loop are Unsloth for training, promptfoo for the gate, and vLLM on Modal for serving.
The Result
The numbers depend on how the model is called, so here are both.
Served in INT8 and decoded greedily, the first production candidate passed 88 of 95 held-out cases (92.6%). Re-runs landed between 89.5% and 92.6%, with 6 cases failing every time and 6 that flipped between runs.
We shipped it anyway, because INT8 was 2.6 times faster and the failures were concentrated in a few hard cases.
The 95 of 95 comes from the current setup: the prompt that describes the MDMA-DSL grammar, temperature 1 and our latest validator. Without it the model misreads the input.
Across the six promptfoo suites the same setup passes 181 of 181 committed cases, and the median generation on the holdout run took about 1.4 seconds.
Serving the Result
We publish the merged model in BF16 and serve it INT8-quantised behind a chat API endpoint on Modal, a managed GPU platform. INT8 roughly halves memory against BF16, and because the endpoint follows the standard chat API format, every harness and SDK we already use works against it unchanged.
Serving needed one workaround that had nothing to do with training. Gemma 4 has a weakness: it can fall into repetitive loops and never produce an answer.
We have not fixed it, only softened it with sampling settings and a loop detector on the client side, and no single setting is airtight. Our evaluated configuration is:
min_p: 0.02
repetition_penalty: 1.1
max output tokens: 2048Our guide to running a self-hosted LLM for business covers when serving your own model pays off.
Matt Sadowski
CEO of Mobile Reality
Own the Model Behind Your AI Feature
We help teams decide whether a smaller, fine-tuned or self-hosted model can take over a narrow, high-volume task, and prove it before anything changes in production.
Fine-tuning open-weight models for narrow tasks with a fixed output format.
A held-out evaluation gate on your own cases before any switch.
Serving behind an OpenAI-compatible endpoint, quantised for your hardware.
Data that stays on infrastructure you control.
An honest answer when a hosted API is still the better choice.
Conclusion
Tuning an open-weight LLM is a precise tool: it changes behaviour, not knowledge, and it pays off on narrow tasks with a fixed output grammar. Our Gemma 4 fine-tune works because the task is narrow, the validator is strict and the holdout and gate existed from the first run.
- Use prompting for instructions, RAG for knowledge that changes, and tuning for behaviour you need thousands of times.
- Start with a small model if it already learns the task, then try a more capable base to see how far you can push it; we went from Gemma 4 E4B to the 26B-A4B.
- LoRA with Unsloth made a 26-billion-parameter mixture-of-experts fine-tune possible on a single A100 80GB GPU.
- Gate the model on held-out cases with a validator, and report the conditions with the score: ours reaches 95 of 95 and 181 of 181 with the grammar prompt at temperature 1, and the first production candidate scored closer to 92%.
- Treat serving settings as part of the model, versioned in one place, because inference failures can look like training failures.
If you have a narrow task that a prompted model handles almost well enough, a fine-tuned open-weight model may close the gap and give you ownership of it. Our custom AI model development services start with the evaluation gate, so you know whether you need the fine-tune before you pay for it.
Frequently Asked Questions
When should I fine-tune an LLM instead of using RAG or prompting?
Use prompting when the model can do the task but needs clearer instructions, and RAG when it lacks knowledge that changes over time. Fine-tune when you need a specific behaviour or output format thousands of times and prompting does not produce it reliably, or when you want a smaller model you own to do the job.
What is the difference between LoRA and QLoRA?
LoRA freezes the base model and trains small additional matrices, which cuts memory use and produces a compact adapter. QLoRA does the same while loading the base model in 4-bit precision, saving more memory at some risk to quality. Our Gemma 4 fine-tune used 16-bit LoRA on a single A100 80GB GPU.
Is fine-tuning the same as model distillation?
No. Distillation trains a smaller student model to imitate a larger teacher model, often across many tasks. Supervised fine-tuning, which is what we did, trains a model on input and output pairs for a specific task. The goal was reliability on one narrow job, not a general copy of a bigger model.
How do you evaluate a fine-tuned model before using it?
Keep a held-out set of cases out of training and check every output automatically. Our gate holds 95 scenarios that must all pass our MDMA validator, and the current model passes all 95, plus 181 of 181 cases across six evaluation suites.
More on AI Cost, Small Models and Evaluation
What an AI system costs, which model it runs on, and how you prove it still works are one decision, not three. These articles cover the build and run cost, model choice and the evaluation that makes switching safe:
- LLM Router: How We Route Models by Cost, Quality and Failure
- Prompt Caching: A Practical Guide for Claude, OpenAI, Gemini
- Self-Hosted LLM Guide: Hardware, Tools and When It Pays Off
- SLM vs LLM: Key Differences and When Small Models Win
- AI Development Cost in 2026: What Drives It and Real Ranges
- LLM Evaluation in Practice: How We Test Across 32 Models
- Context Is King: How to Serve It With Context Engineering
- Structured LLM Output Without JSON Schemas | MDMA
- Build AI Agents with 75+ Deployments Cutting Costs 60% in 2026
Want to know whether a smaller or self-hosted model could handle part of your workload? Our custom AI model development team starts with an evaluation on your own cases.
