Training Data
- In Turkish
- Eğitim Verisi
In short
Training data is the set of examples a machine learning model learns from, and its quality, size, and coverage largely determine how well the model performs.
What is training data?
Training data is the collection of examples used to teach a machine learning model. Depending on the task, it might be emails marked spam or not spam, photos tagged with the objects they show, rows of past sales with their final prices, or, for a large language model, trillions of tokens of text and code. During training, the model adjusts its parameters until its outputs match the patterns in this data.
Before training, a dataset is usually split into three parts. The training set is what the model learns from, the validation set is used to compare settings and decide when to stop, and the test set is kept aside until the end to estimate how the model will do on data it has never seen. Preparing training data often takes more time than training itself, because it has to be collected, cleaned, deduplicated, labeled, and checked for errors.
Think of training data as the textbook and practice problems a student studies from. A student who only saw easy problems, or whose answer key had mistakes, will struggle on the real exam, and a model is the same: gaps, errors, and biases in the training data turn into gaps, errors, and biases in the model. This idea is often summed up as garbage in, garbage out.
Training data is often confused with the input a model receives at run time. The documents you paste into a prompt or retrieve with RAG are context for a single request, not training data, and they don't change the model's weights. It is also different from test data: if test examples leak into the training set, a problem called data leakage, the model's scores look great in evaluation but fall apart in real use.
Key takeaways
- Training data is the set of examples a model learns its patterns from.
- Datasets are usually split into training, validation, and test sets.
- Errors, gaps, and biases in the data carry over into the model's behavior.
- Data in a prompt or retrieved by RAG is context for one request, not training data.
- Keeping test data separate from training data avoids misleading results.
Example
import random
# 1,000 labeled examples: (email text, is_spam)
examples = [(f"email {i}", i % 5 == 0) for i in range(1000)]
random.seed(42)
random.shuffle(examples) # shuffle so each split is representative
train = examples[:800] # 80%: the model learns from these
validation = examples[800:900] # 10%: used to tune settings
test = examples[900:] # 10%: used only for the final score
print(len(train), len(validation), len(test)) # 800 100 100Readers ask
What is the difference between training data and test data?
Training data is what the model learns from, while test data is held back and used only to measure how well the finished model handles examples it has never seen. Using the same examples for both would hide overfitting.
How much training data do you need?
It depends on the task and the model. A simple classifier may work with a few thousand good examples, fine-tuning an LLM often needs hundreds to thousands, and training a large model from scratch needs billions of examples or more; quality and variety usually matter as much as quantity.
What is data leakage?
Data leakage happens when information from the test set, or from the future, sneaks into the training data. The model then looks far more accurate during evaluation than it will be in production.
See also
- Machine LearningAI & Machine Learning, p. 27Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- Supervised LearningAI & Machine Learning, p. 43Supervised learning is machine learning where a model learns from labeled examples, inputs paired with correct answers, to predict outputs for new data.
- Unsupervised LearningAI & Machine Learning, p. 50Unsupervised learning is machine learning in which a model finds patterns, groups, or structure in unlabeled data, without being given the correct answers.
- OverfittingAI & Machine Learning, p. 34Overfitting happens when a machine learning model learns its training data so closely, including its noise, that it performs poorly on new, unseen data.
- Fine-tuningAI & Machine Learning, p. 19Fine-tuning is the process of taking a pretrained machine learning model and training it further on a smaller, specific dataset to adapt it to one task.
- Model ParametersAI & Machine Learning, p. 30Model parameters are the internal numbers, such as weights and biases, that a machine learning model learns in training and uses to turn inputs into outputs.
Spotted a mistake or something missing on this page?Suggest an edit