Skip to main content

Training Data

In Turkish
Eğitim Verisi
Updated 3 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/training-data

In short

Training data is the set of examples a machine learning model learns from, and its quality, size, and coverage largely determine how well the model performs.

What is training data?

Training data is the collection of examples used to teach a machine learning model. Depending on the task, it might be emails marked spam or not spam, photos tagged with the objects they show, rows of past sales with their final prices, or, for a large language model, trillions of tokens of text and code. During training, the model adjusts its parameters until its outputs match the patterns in this data.

Before training, a dataset is usually split into three parts. The training set is what the model learns from, the validation set is used to compare settings and decide when to stop, and the test set is kept aside until the end to estimate how the model will do on data it has never seen. Preparing training data often takes more time than training itself, because it has to be collected, cleaned, deduplicated, labeled, and checked for errors.

Think of training data as the textbook and practice problems a student studies from. A student who only saw easy problems, or whose answer key had mistakes, will struggle on the real exam, and a model is the same: gaps, errors, and biases in the training data turn into gaps, errors, and biases in the model. This idea is often summed up as garbage in, garbage out.

Training data is often confused with the input a model receives at run time. The documents you paste into a prompt or retrieve with RAG are context for a single request, not training data, and they don't change the model's weights. It is also different from test data: if test examples leak into the training set, a problem called data leakage, the model's scores look great in evaluation but fall apart in real use.

Key takeaways

  • Training data is the set of examples a model learns its patterns from.
  • Datasets are usually split into training, validation, and test sets.
  • Errors, gaps, and biases in the data carry over into the model's behavior.
  • Data in a prompt or retrieved by RAG is context for one request, not training data.
  • Keeping test data separate from training data avoids misleading results.

Example

Splitting a dataset into training, validation, and test setspython
import random

# 1,000 labeled examples: (email text, is_spam)
examples = [(f"email {i}", i % 5 == 0) for i in range(1000)]
random.seed(42)
random.shuffle(examples)  # shuffle so each split is representative

train = examples[:800]          # 80%: the model learns from these
validation = examples[800:900]  # 10%: used to tune settings
test = examples[900:]           # 10%: used only for the final score

print(len(train), len(validation), len(test))  # 800 100 100

Readers ask

What is the difference between training data and test data?

Training data is what the model learns from, while test data is held back and used only to measure how well the finished model handles examples it has never seen. Using the same examples for both would hide overfitting.

How much training data do you need?

It depends on the task and the model. A simple classifier may work with a few thousand good examples, fine-tuning an LLM often needs hundreds to thousands, and training a large model from scratch needs billions of examples or more; quality and variety usually matter as much as quantity.

What is data leakage?

Data leakage happens when information from the test set, or from the future, sneaks into the training data. The model then looks far more accurate during evaluation than it will be in production.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings