Deep Learning
- In Turkish
- Derin Öğrenme
In short
Deep learning is a subset of machine learning that uses neural networks with many layers to learn complex patterns from raw data such as images and text.
What is deep learning?
Deep learning is a branch of machine learning based on neural networks with many hidden layers, which is where the word deep comes from. Each layer transforms its input into a slightly more abstract representation, so the network can learn complex patterns step by step. Modern deep learning models range from millions to hundreds of billions of parameters.
A key advantage of deep learning is that it discovers useful features on its own. In classic machine learning, engineers often had to hand-design features, such as counting edges in an image; a deep network given raw pixels learns to detect edges in its early layers, shapes in the middle layers, and whole objects like faces or cars in later layers. Training these large models requires a lot of data and specialized hardware such as GPUs, which can perform many calculations in parallel.
An analogy is learning to read: first you recognize strokes, then letters, then words, and finally the meaning of whole sentences, with each level building on the one before. Deep learning powers image and speech recognition, machine translation, the perception systems of self-driving cars, recommendations, and large language models.
Deep learning is often used interchangeably with AI or machine learning, but it is a subset of both: AI is the broad goal, machine learning is learning from data, and deep learning is machine learning with deep neural networks. For small, structured datasets, such as a spreadsheet of customer records, simpler machine learning methods are often faster, cheaper, and just as accurate.
Key takeaways
- Deep learning uses neural networks with many hidden layers.
- Each layer learns a more abstract representation of the data.
- It learns features automatically from raw data such as pixels, audio, or text.
- It needs large datasets and parallel hardware such as GPUs.
- For small tabular datasets, classic machine learning is often a better fit.
Example
import numpy as np
rng = np.random.default_rng(0)
layer_sizes = [784, 256, 128, 64, 10] # input, three hidden layers, output
# One weight matrix per pair of neighboring layers (random until trained)
weights = [rng.normal(0, 0.1, (a, b)) for a, b in zip(layer_sizes, layer_sizes[1:])]
def forward(x):
for w in weights[:-1]:
x = np.maximum(0, x @ w) # ReLU activation in each hidden layer
return x @ weights[-1] # output layer: one score per class
image = rng.random(784) # a fake 28x28 image, flattened
print(forward(image).shape) # (10,) -> scores for the digits 0-9Readers ask
What is the difference between deep learning and machine learning?
Deep learning is a subset of machine learning that uses neural networks with many layers. Classic machine learning often relies on hand-picked features and works well on smaller, structured data, while deep learning learns features itself and excels at images, audio, and text.
Why is it called deep learning?
The word deep refers to the number of layers in the neural network. A network with many hidden layers between its input and output is considered deep, as opposed to a shallow network with only one or two.
Why does deep learning need GPUs?
Training a deep network involves huge numbers of matrix multiplications, and GPUs can perform thousands of these calculations in parallel. This makes training many times faster than on a regular CPU.
Often compared
See also
- Machine LearningAI & Machine Learning, p. 27Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- Neural NetworkAI & Machine Learning, p. 33A neural network is a machine learning model made of layers of connected artificial neurons that learn patterns from data by adjusting numeric weights.
- TransformerAI & Machine Learning, p. 49A transformer is a neural network architecture that uses attention to weigh how each token in a sequence relates to the others, and it powers most modern LLMs.
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- InferenceAI & Machine Learning, p. 24Inference is the stage where a trained machine learning model is used to make predictions or generate output from new data, without changing what it learned.
Spotted a mistake or something missing on this page?Suggest an edit