LoRA fine-tuning in practice: what works and what does not
LoRA has become the default approach for adapting open-weight models because it is cheaper and faster than full fine-tuning while often reaching comparable quality for narrow tasks. Getting good results from LoRA still requires deliberate choices about rank, target modules, and dataset composition that are easy to get wrong. This post covers the decisions that matter in practice.
Rank selection is a starting point, not a fixed rule
Higher LoRA rank increases the number of trainable parameters in the adapter, which can improve capacity to learn a complex task but also increases the risk of overfitting on small datasets and adds marginally to training and inference cost. A common starting point is a modest rank in the range of eight to thirty-two, adjusted based on validation performance.
Rather than guessing a single rank, run a small sweep across two or three rank values on a held-out validation set before committing to a full training run. The difference in training cost between these trial runs is small compared to the cost of discovering the wrong rank after a full training cycle.
Choosing which modules to target
Applying LoRA only to attention projection matrices is a common default, but including feed-forward layers in the adapted modules can improve performance on tasks that require more than surface-level style adaptation, such as domain-specific reasoning or terminology-heavy generation.
Broadening the set of target modules increases trainable parameters and training cost modestly. For tasks focused on tone, format, or style adaptation, attention-only targeting is often sufficient; for tasks requiring the model to internalize new factual patterns or domain logic, broader targeting tends to perform better.
Dataset size and quality dominate outcomes
LoRA can produce meaningful improvements with dataset sizes far smaller than full fine-tuning requires, sometimes in the low thousands of examples for a narrow task, but quality and consistency of labeling matter more than raw volume. A smaller, carefully curated dataset consistently outperforms a larger, noisy one.
Deduplicate training examples and check for label consistency before training. Inconsistent labels on similar inputs are one of the most common causes of a LoRA adapter that trains without error but performs poorly, since the model has no consistent signal to learn from.
- Prioritize label consistency and deduplication over raw dataset size
- Hold out a validation set that reflects production input distribution, not just training distribution
- Include a small number of intentionally difficult or edge-case examples in the training set
Common failure modes
Overfitting shows up as strong performance on training examples but poor generalization to new inputs, often caused by too few examples, too many training epochs, or a learning rate that is too high for the adapter's small parameter count. Monitor validation loss closely and stop training before it diverges from training loss.
Catastrophic forgetting of general capability, where the adapted model becomes worse at tasks outside the fine-tuning focus, is less severe with LoRA than with full fine-tuning but still occurs, particularly at higher ranks or with training data that is very narrow in scope. Mixing a small proportion of general-purpose examples into the training set helps preserve broader capability.
Evaluating and merging adapters
Evaluate a trained adapter against the same representative test set used to justify the fine-tuning project in the first place, not only against training metrics. A drop in perplexity on held-out fine-tuning data does not guarantee an improvement on the actual downstream task the project set out to solve.
Decide whether to merge the LoRA adapter into the base model weights or serve it separately. Serving multiple adapters against a shared base model is more storage-efficient when supporting several fine-tuned variants, while merging simplifies deployment when only one variant is needed per model instance.
Want this applied to your data?
We scope private AI projects in one call and start with a pilot you can evaluate.