Model Parameters
- In Turkish
- Model Parametreleri
In short
Model parameters are the internal numbers, such as weights and biases, that a machine learning model learns in training and uses to turn inputs into outputs.
What are model parameters?
Parameters are the adjustable numbers inside a machine learning model. In a neural network they are mainly weights, which set how strongly one neuron influences the next, and biases, which shift a neuron's output up or down. Together they encode everything the model has learned: a trained model is essentially its architecture plus a very large file of parameter values.
Parameters start out random and are tuned during training by an algorithm such as gradient descent, which nudges each one to reduce the model's error on the training data. Once training ends, they stay fixed during inference. Their number is a rough measure of a model's size: a 7B model has about 7 billion parameters, and at 16 bits (2 bytes) each, just storing them takes about 14 GB of memory, which is why techniques like quantization matter.
An analogy is a huge mixing desk in a recording studio with billions of knobs. Training slowly turns each knob until the output sounds right, and the final knob positions are the model. More knobs let a model capture more complex patterns, but they also need more data, memory, and computing power, and a well-trained smaller model can beat a poorly trained larger one.
Parameters are often confused with hyperparameters. Hyperparameters are settings people choose before training, such as the learning rate, the number of layers, or the batch size, and the model does not learn them. Both are also different from request settings like temperature or maximum output length, which you pass to an AI API at inference time and which don't change the model at all.
Key takeaways
- Parameters are the learned numbers, mainly weights and biases, inside a model.
- Training adjusts them; inference uses them without changing them.
- Parameter count, such as 7B or 70B, is a rough measure of model size.
- Memory needed is roughly the parameter count times the bytes per parameter.
- Hyperparameters are chosen by people and are not learned from data.
Example
# Count the parameters of a small fully connected neural network
layer_sizes = [784, 128, 64, 10] # input, two hidden layers, output
total = 0
for inputs, outputs in zip(layer_sizes, layer_sizes[1:]):
weights = inputs * outputs # one weight per connection
biases = outputs # one bias per neuron
total += weights + biases
print(total) # 109386
# Memory for a 7-billion-parameter model at 2 bytes per parameter
print(7e9 * 2 / 1e9, "GB") # 14.0 GBReaders ask
What is the difference between parameters and hyperparameters?
Parameters are learned automatically from the training data, while hyperparameters are chosen by people before or during training, such as the learning rate or the number of layers. Hyperparameters control how the parameters get learned.
What does 7B or 70B mean for a model?
It is the approximate number of parameters: 7B means about 7 billion and 70B about 70 billion. Larger models tend to be more capable but need more memory and are slower and more expensive to run.
Do more parameters always make a model better?
No. More parameters give a model more capacity, but results also depend on the amount and quality of training data, the training method, and the task. Smaller, well-trained models often outperform larger ones on specific jobs.
See also
- Neural NetworkAI & Machine Learning, p. 33A neural network is a machine learning model made of layers of connected artificial neurons that learn patterns from data by adjusting numeric weights.
- Gradient DescentAI & Machine Learning, p. 22Gradient descent is an optimization algorithm that trains machine learning models by repeatedly nudging their parameters in the direction that reduces error.
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- QuantizationAI & Machine Learning, p. 37Quantization is a technique that shrinks an AI model by storing its parameters in fewer bits, such as 8 or 4 instead of 16, so inference is faster and cheaper.
- Fine-tuningAI & Machine Learning, p. 19Fine-tuning is the process of taking a pretrained machine learning model and training it further on a smaller, specific dataset to adapt it to one task.
- TemperatureAI & Machine Learning, p. 45Temperature is a setting that controls how random an LLM's output is, from focused and predictable at low values to more varied and creative at high values.
- LoRAAI & Machine Learning, p. 26LoRA is a cheap way to fine-tune a large model: its weights stay frozen and only small added matrices are trained, so a new skill fits in a few megabytes.
Spotted a mistake or something missing on this page?Suggest an edit