Skip to main content

Model Parameters

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/model-parameters

In short

Model parameters are the internal numbers, such as weights and biases, that a machine learning model learns in training and uses to turn inputs into outputs.

What are model parameters?

Parameters are the adjustable numbers inside a machine learning model. In a neural network they are mainly weights, which set how strongly one neuron influences the next, and biases, which shift a neuron's output up or down. Together they encode everything the model has learned: a trained model is essentially its architecture plus a very large file of parameter values.

Parameters start out random and are tuned during training by an algorithm such as gradient descent, which nudges each one to reduce the model's error on the training data. Once training ends, they stay fixed during inference. Their number is a rough measure of a model's size: a 7B model has about 7 billion parameters, and at 16 bits (2 bytes) each, just storing them takes about 14 GB of memory, which is why techniques like quantization matter.

An analogy is a huge mixing desk in a recording studio with billions of knobs. Training slowly turns each knob until the output sounds right, and the final knob positions are the model. More knobs let a model capture more complex patterns, but they also need more data, memory, and computing power, and a well-trained smaller model can beat a poorly trained larger one.

Parameters are often confused with hyperparameters. Hyperparameters are settings people choose before training, such as the learning rate, the number of layers, or the batch size, and the model does not learn them. Both are also different from request settings like temperature or maximum output length, which you pass to an AI API at inference time and which don't change the model at all.

Key takeaways

  • Parameters are the learned numbers, mainly weights and biases, inside a model.
  • Training adjusts them; inference uses them without changing them.
  • Parameter count, such as 7B or 70B, is a rough measure of model size.
  • Memory needed is roughly the parameter count times the bytes per parameter.
  • Hyperparameters are chosen by people and are not learned from data.

Example

Counting the parameters of a small neural networkpython
# Count the parameters of a small fully connected neural network
layer_sizes = [784, 128, 64, 10]  # input, two hidden layers, output

total = 0
for inputs, outputs in zip(layer_sizes, layer_sizes[1:]):
    weights = inputs * outputs  # one weight per connection
    biases = outputs            # one bias per neuron
    total += weights + biases

print(total)  # 109386

# Memory for a 7-billion-parameter model at 2 bytes per parameter
print(7e9 * 2 / 1e9, "GB")  # 14.0 GB

Readers ask

What is the difference between parameters and hyperparameters?

Parameters are learned automatically from the training data, while hyperparameters are chosen by people before or during training, such as the learning rate or the number of layers. Hyperparameters control how the parameters get learned.

What does 7B or 70B mean for a model?

It is the approximate number of parameters: 7B means about 7 billion and 70B about 70 billion. Larger models tend to be more capable but need more memory and are slower and more expensive to run.

Do more parameters always make a model better?

No. More parameters give a model more capacity, but results also depend on the amount and quality of training data, the training method, and the task. Smaller, well-trained models often outperform larger ones on specific jobs.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings