Skip to main content

Inference

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/inference

In short

Inference is the stage where a trained machine learning model is used to make predictions or generate output from new data, without changing what it learned.

What is inference in AI?

In machine learning, inference is the act of running a trained model on new input to get a result. When a photo app labels a picture as a dog, a spam filter flags an email, or an LLM writes a reply to a prompt, the model is performing inference. The model's weights stay fixed during inference; it only applies what it already learned.

The life of a model has two main phases. Training happens once or occasionally, is very expensive, and adjusts the model's weights using large amounts of data. Inference happens every time someone uses the model, often millions of times a day, so its speed and cost matter a lot. For an LLM, inference means generating the answer one token at a time, which is why longer answers take longer and cost more.

An analogy is a chef: training is the years spent learning to cook, while inference is cooking a single dish when an order comes in. Teams measure inference by latency, how long one response takes, and throughput, how many requests the system can handle, and they improve it with techniques like batching requests, caching, and quantization, which stores weights with fewer bits to make a model smaller and faster.

Inference is often confused with training. Chatting with an AI assistant does not train the model: your conversation is part of the input for that request, but the weights don't change unless a separate training or fine-tuning run happens later. Inference can run on powerful servers in the cloud or directly on phones and laptops, which is called on-device or edge inference.

Key takeaways

  • Inference means using a trained model to make predictions on new data.
  • The model's weights do not change during inference.
  • Training happens rarely and costs a lot; inference happens on every request.
  • Latency and throughput are the key inference metrics.
  • For LLMs, inference generates output one token at a time.

Example

Running inference with a trained modelpython
import pickle
import time

# Load a spam classifier that was trained earlier (training is already done)
with open("spam_model.pkl", "rb") as f:
    model = pickle.load(f)

# Inference: run the fixed model on new, unseen input
start = time.perf_counter()
prediction = model.predict(["Congratulations, you won a free prize!"])
latency_ms = (time.perf_counter() - start) * 1000

print(prediction)              # e.g. ['spam']
print(f"{latency_ms:.1f} ms")  # inference latency for one request

Readers ask

What is the difference between training and inference?

Training is the process of teaching a model by adjusting its weights using large amounts of example data. Inference is using the finished model to make predictions on new data, with the weights left unchanged.

Why is AI inference expensive?

Large models perform billions of calculations for every output, and for LLMs this repeats for each generated token. Because inference runs on every request, these costs add up quickly at scale, which is why AI APIs usually charge per token.

What is edge inference?

Edge inference means running a model directly on a local device, such as a phone, laptop, or camera, instead of on a remote server. It reduces latency, works offline, and keeps data on the device, but it requires smaller, optimized models.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings