Unsupervised Learning
- In Turkish
- Denetimsiz Öğrenme
In short
Unsupervised learning is machine learning in which a model finds patterns, groups, or structure in unlabeled data, without being given the correct answers.
What is unsupervised learning?
In unsupervised learning, the training data has no labels: nobody has marked which emails are spam, which customers will leave, or what each photo shows. Instead, the algorithm looks for structure on its own, such as groups of similar items, the main directions in which the data varies, or points that don't fit the usual pattern. It is useful because unlabeled data is plentiful and cheap, while labels often need expensive human work.
The most common tasks are clustering, dimensionality reduction, and anomaly detection. Clustering algorithms such as k-means group similar records together, for example customers with similar buying habits. Dimensionality reduction methods such as principal component analysis (PCA) compress many features into a few that keep most of the information, which helps with visualization and speeds up other models, and anomaly detection flags unusual transactions or server readings that may signal fraud or failure.
Imagine being handed a box of thousands of unlabeled photos and asked to sort them into piles. Nobody tells you the categories, but you notice that some show beaches, some show cities, and some show pets, and you group them accordingly. The hard part, as with any unsupervised result, is deciding whether the piles are meaningful, which is why a person usually has to interpret and name the clusters.
Unsupervised learning is most often contrasted with supervised learning, which trains on labeled examples and learns to predict a known answer, such as a price or a category. It is also different from self-supervised learning, used to pretrain LLMs, where labels are created automatically from the data itself, for example by hiding the next word and asking the model to predict it. Reinforcement learning is a separate approach again, based on rewards rather than labels or structure.
Key takeaways
- Unsupervised learning finds structure in data that has no labels.
- Clustering, dimensionality reduction, and anomaly detection are the main tasks.
- Results need human interpretation, because there is no correct answer to compare against.
- Supervised learning, by contrast, learns from examples paired with correct answers.
- Self-supervised learning creates labels automatically from the data itself.
Example
from sklearn.cluster import KMeans
# Unlabeled customer data: [orders per year, average order value]
customers = [[2, 30], [3, 25], [40, 20], [45, 22], [5, 400], [4, 380]]
# Ask for 3 groups; the algorithm decides which customers belong together
model = KMeans(n_clusters=3, n_init=10, random_state=0)
groups = model.fit_predict(customers)
print(groups) # e.g. [1 1 0 0 2 2]: rare buyers, frequent buyers, big spendersReaders ask
What is the difference between supervised and unsupervised learning?
Supervised learning trains on labeled examples and learns to predict the correct answer for new inputs. Unsupervised learning works with unlabeled data and discovers structure, such as clusters or anomalies, without any answers to learn from.
What are examples of unsupervised learning?
Common examples are customer segmentation, grouping similar documents or news articles, detecting fraudulent transactions or failing machines as anomalies, and compressing data for visualization.
Is LLM pretraining unsupervised learning?
It is usually described as self-supervised learning. The model trains on raw text without human labels, but the training signal comes from the text itself, since the model learns to predict the next token and can check its guess against the real text.
Often compared
See also
- Supervised LearningAI & Machine Learning, p. 43Supervised learning is machine learning where a model learns from labeled examples, inputs paired with correct answers, to predict outputs for new data.
- Machine LearningAI & Machine Learning, p. 27Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- Reinforcement LearningAI & Machine Learning, p. 40Reinforcement learning is a type of machine learning in which an agent learns to make decisions by trial and error, earning rewards for good actions.
- Training DataAI & Machine Learning, p. 48Training data is the set of examples a machine learning model learns from, and its quality, size, and coverage largely determine how well the model performs.
- EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- Semantic SearchAI & Machine Learning, p. 42Semantic search is a search technique that finds results by meaning rather than exact keywords, usually by comparing embeddings of the query and the documents.
Spotted a mistake or something missing on this page?Suggest an edit