LoRA
Low-Rank Adaptation
- Pronunciation
- LOR-uh
In short
LoRA is a cheap way to fine-tune a large model: its weights stay frozen and only small added matrices are trained, so a new skill fits in a few megabytes.
What is LoRA?
Fine-tuning a large language model normally updates all of its billions of weights, which needs a lot of GPU memory and produces a full copy of the model for every task. LoRA, introduced by researchers at Microsoft in 2021, keeps the original weights frozen and adds a pair of small, low-rank matrices next to some layers; only those are trained.
"Low-rank" is what makes it cheap. Instead of learning a full update to a large weight matrix, LoRA learns two thin matrices whose product approximates that update. In the original paper this cut the number of trainable parameters by about 10,000 times for GPT-3. The trained adapter is tiny, can be loaded on top of the base model when needed, and several adapters can share one base model.
LoRA is the most common form of parameter-efficient fine-tuning (PEFT). A popular variant, QLoRA, also stores the frozen base model in 4-bit precision, so fairly large models can be fine-tuned on a single GPU. It is widely used to adapt open models to a domain, a writing style or an output format, and to customize image generation models.
Key takeaways
- LoRA fine-tunes a model by training small added matrices while the original weights stay frozen.
- It needs far less memory than full fine-tuning and produces a small adapter file.
- Adapters can be swapped on top of one shared base model.
- QLoRA combines LoRA with a 4-bit base model to fit training on a single GPU.
Example
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.2-1B")
config = LoraConfig(
r=8, # rank of the added matrices
lora_alpha=16,
target_modules=["q_proj", "v_proj"], # attention layers to adapt
lora_dropout=0.05,
task_type="CAUSAL_LM",
)
model = get_peft_model(model, config)
model.print_trainable_parameters() # only a small share of the weights will trainReaders ask
Is LoRA as good as full fine-tuning?
For most adaptation tasks it comes close, at a fraction of the cost. Full fine-tuning can still do better when the model must learn a lot of genuinely new knowledge, but for style, format and domain adaptation LoRA is usually enough.
What is the difference between LoRA and RAG?
LoRA changes how the model behaves by training a small adapter, while RAG leaves the model unchanged and gives it relevant documents at question time. RAG suits facts that change often; LoRA suits teaching a consistent style, format or skill.
See also
- Fine-tuningAI & Machine Learning, p. 19Fine-tuning is the process of taking a pretrained machine learning model and training it further on a smaller, specific dataset to adapt it to one task.
- QuantizationAI & Machine Learning, p. 37Quantization is a technique that shrinks an AI model by storing its parameters in fewer bits, such as 8 or 4 instead of 16, so inference is faster and cheaper.
- Model ParametersAI & Machine Learning, p. 30Model parameters are the internal numbers, such as weights and biases, that a machine learning model learns in training and uses to turn inputs into outputs.
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- TransformerAI & Machine Learning, p. 49A transformer is a neural network architecture that uses attention to weigh how each token in a sequence relates to the others, and it powers most modern LLMs.
- Training DataAI & Machine Learning, p. 48Training data is the set of examples a machine learning model learns from, and its quality, size, and coverage largely determine how well the model performs.
- RAGAI & Machine Learning, p. 38RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
- Deep LearningAI & Machine Learning, p. 14Deep learning is a subset of machine learning that uses neural networks with many layers to learn complex patterns from raw data such as images and text.
Sources
Spotted a mistake or something missing on this page?Suggest an edit