Skip to main content

LoRA

Low-Rank Adaptation

Pronunciation
LOR-uh
Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/lora

In short

LoRA is a cheap way to fine-tune a large model: its weights stay frozen and only small added matrices are trained, so a new skill fits in a few megabytes.

What is LoRA?

Fine-tuning a large language model normally updates all of its billions of weights, which needs a lot of GPU memory and produces a full copy of the model for every task. LoRA, introduced by researchers at Microsoft in 2021, keeps the original weights frozen and adds a pair of small, low-rank matrices next to some layers; only those are trained.

"Low-rank" is what makes it cheap. Instead of learning a full update to a large weight matrix, LoRA learns two thin matrices whose product approximates that update. In the original paper this cut the number of trainable parameters by about 10,000 times for GPT-3. The trained adapter is tiny, can be loaded on top of the base model when needed, and several adapters can share one base model.

LoRA is the most common form of parameter-efficient fine-tuning (PEFT). A popular variant, QLoRA, also stores the frozen base model in 4-bit precision, so fairly large models can be fine-tuned on a single GPU. It is widely used to adapt open models to a domain, a writing style or an output format, and to customize image generation models.

Key takeaways

  • LoRA fine-tunes a model by training small added matrices while the original weights stay frozen.
  • It needs far less memory than full fine-tuning and produces a small adapter file.
  • Adapters can be swapped on top of one shared base model.
  • QLoRA combines LoRA with a 4-bit base model to fit training on a single GPU.

Example

Adding LoRA adapters with Hugging Face PEFTpython
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.2-1B")
config = LoraConfig(
    r=8,                                  # rank of the added matrices
    lora_alpha=16,
    target_modules=["q_proj", "v_proj"],  # attention layers to adapt
    lora_dropout=0.05,
    task_type="CAUSAL_LM",
)
model = get_peft_model(model, config)
model.print_trainable_parameters()  # only a small share of the weights will train

Readers ask

Is LoRA as good as full fine-tuning?

For most adaptation tasks it comes close, at a fraction of the cost. Full fine-tuning can still do better when the model must learn a lot of genuinely new knowledge, but for style, format and domain adaptation LoRA is usually enough.

What is the difference between LoRA and RAG?

LoRA changes how the model behaves by training a small adapter, while RAG leaves the model unchanged and gives it relevant documents at question time. RAG suits facts that change often; LoRA suits teaching a consistent style, format or skill.

See also

Sources

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings