Skip to main content

Backpropagation

In Turkish
Geri Yayılım
Pronunciation
BAK-prop-uh-GAY-shun
Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/backpropagation

In short

Backpropagation is the algorithm that trains neural networks by measuring how much each weight added to the error and nudging every weight to reduce it.

What is backpropagation?

Training a network means repeating two passes. In the forward pass, an input flows through the layers and produces a prediction, which a loss function compares with the right answer to get an error. In the backward pass, backpropagation works out, for every weight in the network, how the error would change if that weight changed slightly: the gradient.

It does this efficiently with the chain rule from calculus. Starting at the output, it passes the error backwards layer by layer, reusing each layer's result to compute the one before it, so all the gradients come out of a single backward sweep instead of one calculation per weight. Gradient descent then nudges each weight against its gradient.

The method was described several times, and it became widely known through a 1986 paper by David Rumelhart, Geoffrey Hinton and Ronald Williams, which showed that it lets multi-layer networks learn useful internal features. Today frameworks such as PyTorch and TensorFlow do it automatically with automatic differentiation, so developers rarely write gradients by hand.

A common misconception is that backpropagation is the whole learning algorithm. It only computes the gradients; an optimizer such as gradient descent or Adam decides how to use them to update the weights. Problems such as vanishing gradients in very deep networks led to fixes like ReLU activations, careful initialization and residual connections.

Key takeaways

  • Backpropagation computes how much each weight contributed to the error.
  • It uses the chain rule to pass the error backwards layer by layer.
  • One backward pass gives the gradients for all weights at once.
  • An optimizer such as gradient descent then updates the weights.
  • Frameworks like PyTorch do it automatically with autograd.

Example

One training step in PyTorchpython
import torch

model = torch.nn.Sequential(torch.nn.Linear(3, 8), torch.nn.ReLU(), torch.nn.Linear(8, 1))
optimizer = torch.optim.SGD(model.parameters(), lr=0.01)

x = torch.tensor([[0.5, 1.0, -0.2]])
target = torch.tensor([[1.0]])

prediction = model(x)                                   # forward pass
loss = torch.nn.functional.mse_loss(prediction, target)
loss.backward()                                         # backpropagation: fills .grad for every weight
optimizer.step()                                        # gradient descent uses those gradients
optimizer.zero_grad()

Readers ask

What is the difference between backpropagation and gradient descent?

Backpropagation calculates the gradients, how the error changes with each weight. Gradient descent uses those gradients to update the weights. Training needs both.

What is the vanishing gradient problem?

In deep networks, gradients can shrink as they pass backwards through many layers, so the early layers barely learn. ReLU activations, normalization and residual connections reduce the problem.

Do I need to implement backpropagation myself?

Rarely. PyTorch, TensorFlow and JAX compute gradients automatically. Implementing it once for a small network is still a good way to understand how training works.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings