Quick answer
A small open model fine-tuned with LoRA beats a big prompted model when the task is narrow, repetitive and high volume: the same classification thousands of times a day, a fixed set of classes, a strictly defined output format. You get lower latency and cost per request, a shorter prompt, stable JSON, and the option to run the model on your own server. A big model with a prompt is the better choice when you have few examples, the class definitions change often, or the task needs broad knowledge. One test settles it: run both options on the same held-out dataset that neither of them has seen.
The example we will walk through: a query classifier
Take a typical customer service task. Messages arrive from a web form, email and chat. Each one has to be assigned to one of eight categories (complaint, billing, account change, outage, sales question, cancellation, service complaint, other) and given a priority. The result goes into a ticketing system, so it has to be machine readable:
{"category": "billing", "priority": "normal", "confidence": 0.93}
Option A is a big model behind an API with a prompt that describes the categories and gives a few examples. Option B is an open model in the 1-8B parameter class, fine-tuned with a LoRA adapter on historical tickets with correct labels. Below we go through the decision and the build of option B step by step.
When a small model with LoRA wins
Latency. Token generation in language models is mostly limited by memory bandwidth: for every new token the GPU has to read the model weights. A model several times smaller reads several times less data per token. On top of that, a fine-tuned model does not need a long prompt with class definitions and examples, because it learned them from the data. A shorter prompt means less input processing before the first output token.
Cost per request. With an API you pay for input and output tokens. A prompt that describes eight categories with a dozen examples is often longer than the message being classified, and you pay for it on every request. A small model on your own GPU has a mostly fixed cost (hardware, power, upkeep), so it pays off at high, predictable volume.
Stable format. A model fine-tuned on thousands of "message, correct JSON" pairs very rarely renames a field or invents a new category. It is worth knowing that format stability can also be enforced without fine-tuning: vLLM has structured outputs that constrain decoding to a list of values (choice), a regular expression or a JSON schema. Fine-tuning improves how accurately the class is chosen, and schema-constrained decoding guarantees valid syntax. In production, use both.
Privacy and control. Ticket text often contains personal data. A small model runs on your server or in your cloud account, so the data does not go to an outside provider. You also control the version: the model does not change unless you decide it should, and last month's result can be reproduced.
There is research behind this too. In the LoRA Land report (Predibase, 2024), 310 models fine-tuned with 4-bit LoRA across 31 tasks scored on average 34 points higher than their base models and 10 points higher than GPT-4. Bucher and Martini (2024) showed that smaller fine-tuned models consistently and significantly outperform large zero-shot models, including GPT-4 and Claude Opus, in text classification (sentiment, emotion, stances). These are results on specific tasks, not a guarantee for yours, so you still have to measure on your own data.
When fine-tuning does not pay off
You have a few dozen examples. Le Scao and Rush (2021) estimated that in classification a good prompt is worth hundreds of training examples on average. With little data, a big model with a good prompt will usually win. OpenAI's fine-tuning docs say it plainly: if 50 good examples do not help, rethink the task or the prompt before adding more data.
The task changes often. If a new category appears every week or an existing definition shifts, every change means new labels, training and evaluation. You can change a prompt in a minute.
You need broad knowledge. A small model knows less about the world. If correct classification depends on understanding regulations, many industries or rare product names, the big model has the edge. Knowledge that changes, such as the current price list, is better supplied in the context (RAG) than trained into the weights.
Volume is low. At a few hundred requests a day the API bill is small, and your own model means upkeep: a server, monitoring, updates, retraining. We cover that trade-off in more detail in should you build or buy AI.
The comparison in one table
| Criterion | Big model with a prompt | Small model + LoRA |
|---|---|---|
| Getting started | hours: a prompt and a few examples | days or weeks: labels, training, evaluation |
| Data needed | a handful to a dozen examples in the prompt | hundreds to thousands of labelled examples |
| Latency | higher: big model and long prompt | lower: small model and short prompt |
| Cost per request | grows linearly with volume and prompt length | mostly fixed GPU cost, low marginal cost |
| Format stability | good with a schema, depends on model version | very good, version under your control |
| Changing class definitions | edit the prompt | new labels and retraining |
| Broad knowledge | high | limited to what the base model has |
| Data leaves the company | yes, unless the model is hosted by you | no, the model runs on your infrastructure |
Step 1: data, meaning labels, quality and a test set
How many examples. OpenAI recommends starting with 50 well-crafted examples and says improvements usually show up at 50-100. The LIMA authors (2023) tuned a model into assistant behaviour with 1,000 carefully curated examples. For a classifier it makes more sense to count per class: start with a few dozen examples for every category, then add more where the confusion matrix shows errors. Rare classes (for example, service complaints) need deliberate sampling, because a random sample will contain too few of them.
Quality over quantity. The model will learn your inconsistencies. If two people label similar tickets differently, the model will guess. Before training:
- write down a definition of every class with borderline examples;
- give the same sample to two people and check agreement; where they differ, fix the definitions;
- remove duplicates and near duplicates (the same email templates), because they inflate the score;
- check that the examples look like what the model will see in production: the same channels, lengths and typos.
A held-out test set. Before you start training, set aside 10-20% of the data and do not look at it while tuning. Ideally hold out the most recent data (for example, the last month), because that tells you how the model will handle future tickets. Near duplicates must land entirely on one side of the split. Carve out a separate small validation set for choosing hyperparameters and for calibration. Synthetic data generated by a big model can fill out rare classes in training, but it should never go into the test set.
Step 2: a baseline with the big model
Before you train anything, measure option A on the same test set. Write the best prompt you can, with class definitions and examples, and turn on JSON schema decoding. That is the bar. Fine-tuning only makes sense if the small model clears it, or comes close at clearly lower cost and latency. Without this baseline you do not know whether training achieved anything.
Step 3: LoRA, QLoRA or full fine-tuning
LoRA (Hu et al., 2021) freezes the model weights and trains only small low-rank matrices added to selected layers. The authors report that for GPT-3 175B it cuts the number of trainable parameters by 10,000 times and GPU memory needs by 3 times, with no extra inference latency once the weights are merged. QLoRA (Dettmers et al., 2023) goes further: it keeps the frozen base model in 4 bits (NF4), which made it possible to fine-tune a 65B model on a single 48 GB GPU.
| Full fine-tuning | LoRA | QLoRA | |
|---|---|---|---|
| What is trained | all weights | low-rank adapters, base frozen in BF16 | adapters, base frozen in 4 bits |
| Memory for weights and optimizer | about 16 bytes per parameter with mixed-precision Adam (about 64 GB for a 4B model, without activations) | about 2 bytes per base parameter plus a small adapter (about 8 GB for 4B) | about 0.5 bytes per base parameter plus the adapter (about 2-3 GB for 4B) |
| Time | longest, often many GPUs | short, usually one GPU | similar to LoRA, with dequantization overhead |
| Quality on a narrow task | the reference | usually comparable when LoRA covers all layers | close to LoRA |
| Artifact | a full copy of the model | an adapter in the megabytes to hundreds of MB range | an adapter |
The memory figures cover weights and optimizer state, without activations, which grow with sequence length and batch size. The 16 bytes per parameter for mixed-precision Adam comes from the ZeRO analysis (Rajbhandari et al., 2019).
What recent research says about quality:
- LoRA on all layers. "LoRA Without Regret" (Thinking Machines, 2025) shows that LoRA applied to the MLP layers as well learns with the same sample efficiency as full fine-tuning until the dataset exceeds the adapter's capacity. Attention-only LoRA underperforms, even with the same number of trainable parameters. In the PEFT library this corresponds to
target_modules="all-linear". - Learning rate. The same analysis finds that the optimal learning rate for LoRA is about 10 times higher than for full fine-tuning.
- Learns less, forgets less. Biderman et al. (2024) showed that on large datasets of new knowledge (code, math) LoRA clearly trails full fine-tuning, but it better preserves what the model could already do.
The takeaway for a classifier: the task is narrow and the data runs from hundreds to a few thousand examples, so LoRA on all linear layers is a sensible default. Choose QLoRA when the model does not fit in GPU memory in BF16. Full fine-tuning is rarely needed here.
Step 4: evaluation, meaning more than accuracy
Accuracy and macro-F1. Accuracy alone misleads when classes are imbalanced: if 60% of tickets are billing questions, a model that always answers "billing" scores 60%. Macro-F1 computes F1 separately for each class and averages them, so a weak rare class pulls the score down.
The confusion matrix. Rows are true classes, columns are predictions. It shows which pairs of classes the model confuses, for example "complaint" and "service complaint". It is the best pointer to where definitions need fixing or more examples are needed. Often it turns out that people confuse the same pair, and then the problem is the taxonomy, not the model.
Calibration. If the model returns a confidence, it has to mean something. Guo et al. (2017) showed that modern neural networks are often overconfident. You check this with a reliability diagram: split predictions into confidence bins and compute the hit rate in each. Among predictions with confidence around 0.8, about 80% should be correct. You can read confidence from the token probabilities (logprobs) of the class label instead of asking the model to write a number. If calibration is poor, temperature scaling tuned on the validation set is a simple and effective fix.
A threshold and a fallback path. Good calibration lets you set a threshold: below it, the ticket goes to a person or to the big model. On the test set, compute what share of requests passes automatically and what accuracy you get on that share. That usually convinces the business more than a single accuracy number.
| Metric | What it tells you | When it misleads |
|---|---|---|
| Accuracy | share of correct answers | with imbalanced classes |
| Macro-F1 | average quality across all classes | when classes differ a lot in business weight |
| Confusion matrix | which classes get mixed up | with a small test set the counts are too sparse |
| Reliability diagram, ECE | whether confidence matches accuracy | with a few dozen examples the bins are too empty |
Step 5: serving the adapters
After training you have two paths. You can merge the adapter into the base model (in PEFT, merge_and_unload()) and serve it like a normal model, with no adapter overhead. Or you can keep the adapters separate and serve many of them on one base model.
The second path makes sense when you have several classifiers, for example one per department or per language. vLLM supports this natively:
vllm serve <base-model> \
--enable-lora \
--lora-modules routing-pl=/adapters/routing-pl routing-en=/adapters/routing-en \
--max-loras 4 \
--max-lora-rank 16
The client picks an adapter by name in the model field, just as it picks a model in an OpenAI-compatible API. Requests for different adapters and for the base model share the same batch. Set --max-lora-rank to the highest rank among your adapters, because a higher value wastes memory. Adapters can also be loaded at runtime (/v1/load_lora_adapter after setting VLLM_ALLOW_RUNTIME_LORA_UPDATING), but the vLLM docs recommend this only in an isolated, fully trusted environment.
The scale of this approach is well studied. S-LoRA (Sheng et al., 2023) serves thousands of adapters on a single GPU by keeping them in main memory and moving only the ones currently needed to the GPU. In LoRA Land, 25 fine-tuned Mistral-7B models ran on a single A100 80 GB. At pimento we do this on NVIDIA Blackwell, and our approach to models and fine-tuning is described on the pillar page.
When serving the classifier, add schema-constrained decoding (structured_outputs with a choice list or a JSON schema). Then even a rare model error will not break the parser in the ticketing system.
Decision checklist
Fine-tuning a small model with LoRA makes sense if you answer "yes" to most of these:
- The task is narrow and has a fixed set of outputs (classes, fields, format).
- Volume is at least thousands of requests a day, or latency matters to the user.
- You have, or can get, hundreds of well-labelled examples, rare classes included.
- Class definitions stay stable for months, not weeks.
- The data should not leave your infrastructure, or you want to control the model version.
- You have a held-out test set and a measured baseline from the big model.
- Someone will maintain the model: quality monitoring and retraining when the data drifts.
If you answer "no" to questions 3, 4 or 6, start with a big model, a good prompt and a JSON schema. Meanwhile, collect corrected labels from production, because they will become your fine-tuning data once volume grows.
Sources
- Hu et al.: LoRA, Low-Rank Adaptation of Large Language Models (2021)
- Dettmers et al.: QLoRA, Efficient Finetuning of Quantized LLMs (2023)
- Schulman et al., Thinking Machines: LoRA Without Regret (2025)
- Biderman et al.: LoRA Learns Less and Forgets Less (2024)
- Zhao et al.: LoRA Land, 310 Fine-tuned LLMs that Rival GPT-4 (2024)
- Bucher, Martini: Fine-Tuned "Small" LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification (2024)
- Le Scao, Rush: How Many Data Points is a Prompt Worth? (2021)
- Zhou et al.: LIMA, Less Is More for Alignment (2023)
- Rajbhandari et al.: ZeRO, Memory Optimizations Toward Training Trillion Parameter Models (2019)
- Guo et al.: On Calibration of Modern Neural Networks (2017)
- Sheng et al.: S-LoRA, Serving Thousands of Concurrent LoRA Adapters (2023)
- OpenAI: Supervised fine-tuning
- Hugging Face PEFT: LoRA
- vLLM: LoRA Adapters
- vLLM: Structured Outputs
- scikit-learn: Probability calibration
- NVIDIA: Mastering LLM Techniques, Inference Optimization
