Skip to main content

Supervised Learning

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/supervised-learning

In short

Supervised learning is machine learning where a model learns from labeled examples, inputs paired with correct answers, to predict outputs for new data.

What is supervised learning?

In supervised learning, every training example comes with a label, the answer the model should produce. A spam filter learns from emails marked spam or not spam, and a house-price model learns from past sales with their final prices. During training, the model makes predictions, compares them with the labels using a loss function that measures how wrong it was, and adjusts its internal parameters to reduce that error.

Supervised tasks fall into two main groups. Classification predicts a category, such as whether a photo shows a cat or a dog, or whether a transaction is fraudulent. Regression predicts a number, such as tomorrow's temperature or a delivery time. Models are always evaluated on a held-out test set they never saw during training, to check that they learned general patterns instead of memorizing the examples.

It works like studying with flashcards that have the answer on the back: you guess, flip the card, and correct yourself until you get new cards right. Supervised learning powers much of everyday AI, including image recognition, speech-to-text, medical image analysis, credit scoring, and recommendation ranking. Its main cost is labeling, because collecting thousands of correct answers often requires human experts.

Supervised learning is usually contrasted with unsupervised learning, which works with unlabeled data and looks for structure on its own, such as grouping customers into clusters or spotting unusual behavior. Reinforcement learning is a third approach, where an agent learns from rewards rather than correct answers. Large language models are first pretrained with self-supervised learning, where the labels come from the text itself by predicting the next token, and are then often refined with supervised fine-tuning on example answers.

Key takeaways

  • Supervised learning trains on labeled data: inputs paired with the correct outputs.
  • Classification predicts categories; regression predicts numbers.
  • A loss function measures errors, and training adjusts the model to reduce them.
  • Unsupervised learning finds patterns in unlabeled data instead.
  • Collecting high-quality labels is often the most expensive part.

Example

Training a classifier on labeled data (scikit-learn)python
from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

# Labeled data: flower measurements (X) and the correct species (y)
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)  # learn from inputs paired with answers

print(model.score(X_test, y_test))  # accuracy on flowers it has never seen
print(model.predict(X_test[:3]))    # predicted species for new examples

Readers ask

What is the difference between supervised and unsupervised learning?

Supervised learning trains on labeled data, where each example includes the correct answer, and learns to predict that answer. Unsupervised learning trains on unlabeled data and discovers structure by itself, such as clusters of similar items.

What are examples of supervised learning?

Common examples are spam detection, image classification, speech recognition, fraud detection, and predicting prices or demand. In each case the model learns from past examples where the right answer is already known.

What is semi-supervised learning?

Semi-supervised learning combines a small amount of labeled data with a large amount of unlabeled data. It is useful when labels are expensive, because the unlabeled examples help the model learn the overall structure of the data.

Often compared

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings