Gradient Descent
- In Turkish
- Gradyan İnişi
In short
Gradient descent is an optimization algorithm that trains machine learning models by repeatedly nudging their parameters in the direction that reduces error.
What is gradient descent?
Gradient descent is the algorithm most machine learning models use to learn. Training starts with a model whose parameters, the internal numbers it uses to make predictions, are set to random or default values. A loss function measures how wrong the model's predictions are on the training data, and gradient descent's job is to find parameter values that make that loss as small as possible.
At each step, the algorithm calculates the gradient, which says how much the loss would change if each parameter were increased slightly. It then moves every parameter a small amount in the opposite direction, downhill, and repeats. The size of each step is set by the learning rate: too large and training overshoots and becomes unstable, too small and it crawls. Because computing the gradient over millions of examples is slow, most training uses stochastic or mini-batch gradient descent, which estimates the gradient from a small random batch of examples at a time.
The classic analogy is walking down a mountain in thick fog. You can't see the valley, but you can feel which way the ground slopes under your feet, so you take a step downhill, check again, and repeat until the ground feels flat. Gradient descent and its variants, such as the widely used Adam optimizer, train almost everything from linear regression to deep neural networks and large language models.
Gradient descent is often confused with backpropagation. Backpropagation is the method that efficiently calculates the gradients in a neural network, working backward from the output layer, while gradient descent is the rule that uses those gradients to update the weights; training needs both. Gradient descent also doesn't guarantee the best possible solution, since it can settle in a local minimum or a flat region, but in large neural networks the solutions it finds are usually good enough.
Key takeaways
- Gradient descent minimizes a loss function by adjusting a model's parameters step by step.
- Each step moves the parameters opposite to the gradient, the direction in which the error grows fastest.
- The learning rate sets the step size and strongly affects whether training succeeds.
- Mini-batch gradient descent estimates the gradient from small random batches of data.
- Backpropagation computes the gradients; gradient descent uses them to update the weights.
Example
# Learn the weight w in y = w * x from examples where the true answer is w = 2
xs = [1.0, 2.0, 3.0, 4.0]
ys = [2.0, 4.0, 6.0, 8.0]
w = 0.0 # start with a poor guess
learning_rate = 0.01
for step in range(200):
# Gradient of the mean squared error with respect to w
grad = sum(2 * (w * x - y) * x for x, y in zip(xs, ys)) / len(xs)
w -= learning_rate * grad # take a small step downhill
print(round(w, 3)) # 2.0Readers ask
What is the learning rate in gradient descent?
The learning rate is a number that controls how big a step the algorithm takes on each update. If it is too high, the loss can bounce around or explode; if it is too low, training takes far too long, so it is one of the most important settings to tune.
What is the difference between gradient descent and backpropagation?
Backpropagation calculates the gradient of the loss with respect to every weight in a neural network. Gradient descent then uses those gradients to update the weights, so backpropagation works out which way to move and gradient descent actually moves.
What is stochastic gradient descent?
Stochastic gradient descent (SGD) updates the parameters using the gradient from a single example or a small random batch instead of the entire dataset. Each step is noisier but much cheaper, which makes training on large datasets practical.
See also
- Machine LearningAI & Machine Learning, p. 27Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- Neural NetworkAI & Machine Learning, p. 33A neural network is a machine learning model made of layers of connected artificial neurons that learn patterns from data by adjusting numeric weights.
- Model ParametersAI & Machine Learning, p. 30Model parameters are the internal numbers, such as weights and biases, that a machine learning model learns in training and uses to turn inputs into outputs.
- Training DataAI & Machine Learning, p. 48Training data is the set of examples a machine learning model learns from, and its quality, size, and coverage largely determine how well the model performs.
- Deep LearningAI & Machine Learning, p. 14Deep learning is a subset of machine learning that uses neural networks with many layers to learn complex patterns from raw data such as images and text.
- OverfittingAI & Machine Learning, p. 34Overfitting happens when a machine learning model learns its training data so closely, including its noise, that it performs poorly on new, unseen data.
- BackpropagationAI & Machine Learning, p. 6Backpropagation is the algorithm that trains neural networks by measuring how much each weight added to the error and nudging every weight to reduce it.
Spotted a mistake or something missing on this page?Suggest an edit