Skip to main content

Gradient Descent

In Turkish
Gradyan İnişi
Updated 3 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/gradient-descent

In short

Gradient descent is an optimization algorithm that trains machine learning models by repeatedly nudging their parameters in the direction that reduces error.

What is gradient descent?

Gradient descent is the algorithm most machine learning models use to learn. Training starts with a model whose parameters, the internal numbers it uses to make predictions, are set to random or default values. A loss function measures how wrong the model's predictions are on the training data, and gradient descent's job is to find parameter values that make that loss as small as possible.

At each step, the algorithm calculates the gradient, which says how much the loss would change if each parameter were increased slightly. It then moves every parameter a small amount in the opposite direction, downhill, and repeats. The size of each step is set by the learning rate: too large and training overshoots and becomes unstable, too small and it crawls. Because computing the gradient over millions of examples is slow, most training uses stochastic or mini-batch gradient descent, which estimates the gradient from a small random batch of examples at a time.

The classic analogy is walking down a mountain in thick fog. You can't see the valley, but you can feel which way the ground slopes under your feet, so you take a step downhill, check again, and repeat until the ground feels flat. Gradient descent and its variants, such as the widely used Adam optimizer, train almost everything from linear regression to deep neural networks and large language models.

Gradient descent is often confused with backpropagation. Backpropagation is the method that efficiently calculates the gradients in a neural network, working backward from the output layer, while gradient descent is the rule that uses those gradients to update the weights; training needs both. Gradient descent also doesn't guarantee the best possible solution, since it can settle in a local minimum or a flat region, but in large neural networks the solutions it finds are usually good enough.

Key takeaways

  • Gradient descent minimizes a loss function by adjusting a model's parameters step by step.
  • Each step moves the parameters opposite to the gradient, the direction in which the error grows fastest.
  • The learning rate sets the step size and strongly affects whether training succeeds.
  • Mini-batch gradient descent estimates the gradient from small random batches of data.
  • Backpropagation computes the gradients; gradient descent uses them to update the weights.

Example

Learning one weight with gradient descentpython
# Learn the weight w in y = w * x from examples where the true answer is w = 2
xs = [1.0, 2.0, 3.0, 4.0]
ys = [2.0, 4.0, 6.0, 8.0]

w = 0.0              # start with a poor guess
learning_rate = 0.01

for step in range(200):
    # Gradient of the mean squared error with respect to w
    grad = sum(2 * (w * x - y) * x for x, y in zip(xs, ys)) / len(xs)
    w -= learning_rate * grad  # take a small step downhill

print(round(w, 3))  # 2.0

Readers ask

What is the learning rate in gradient descent?

The learning rate is a number that controls how big a step the algorithm takes on each update. If it is too high, the loss can bounce around or explode; if it is too low, training takes far too long, so it is one of the most important settings to tune.

What is the difference between gradient descent and backpropagation?

Backpropagation calculates the gradient of the loss with respect to every weight in a neural network. Gradient descent then uses those gradients to update the weights, so backpropagation works out which way to move and gradient descent actually moves.

What is stochastic gradient descent?

Stochastic gradient descent (SGD) updates the parameters using the gradient from a single example or a small random batch instead of the entire dataset. Each step is noisier but much cheaper, which makes training on large datasets practical.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings