Skip to main content

Token

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/llm-token

In short

A token is the basic unit of text that an LLM reads and generates, usually a whole word, part of a word, or a punctuation mark, mapped to a numeric ID.

What is a token in an LLM?

Language models don't read text letter by letter or word by word; they read tokens. A token is a small chunk of text, such as a common word like 'the', a piece of a longer word like 'ing', a punctuation mark, or a space followed by a word. Before any text reaches the model, a component called a tokenizer splits it into tokens and converts each one into a number, its token ID.

Most tokenizers use subword methods such as byte-pair encoding, which learn a vocabulary of frequent chunks from training text. Common words become a single token, while rare words, names, and code are split into several pieces. In English, one token averages about four characters, or roughly three quarters of a word, but other languages, emoji, and source code often need more tokens for the same amount of text.

Tokens matter to developers because almost everything is measured in them. Context windows, output limits, rate limits, and API pricing are all counted in tokens, and generation speed is often quoted in tokens per second. A handy analogy is a mobile data plan measured in gigabytes rather than in web pages: the token is the unit an LLM uses for capacity and billing.

An LLM token is not the same as a security token, such as a JWT or an API key, which proves identity or permission. It is also different from an embedding: a token is a piece of text with an ID, and inside the model each token ID is turned into an embedding vector that carries its meaning.

Key takeaways

  • A token is a chunk of text: a word, part of a word, or a symbol.
  • A tokenizer splits text into tokens and maps each one to a numeric ID.
  • In English, one token is roughly four characters or three quarters of a word.
  • Context windows, limits, and pricing are all measured in tokens.
  • Different models use different tokenizers, so token counts vary between them.

Example

How text is split into tokenstypescript
// One way a tokenizer might split a sentence (the exact split varies by model)
const text = "Tokenization isn't hard!";
const tokens = ["Token", "ization", " isn", "'t", " hard", "!"];

console.log(tokens.join("") === text); // true: the tokens rebuild the original text
console.log(tokens.length);            // 6 tokens

// Rule of thumb for English text: about 4 characters per token
function estimateTokens(input: string): number {
  return Math.ceil(input.length / 4);
}

console.log(estimateTokens(text)); // 6 (24 characters / 4)

Readers ask

How many words is 1,000 tokens?

For typical English text, 1,000 tokens is roughly 750 words. The exact number depends on the tokenizer and the text, and code or non-English languages usually use more tokens per word.

Why do LLMs use tokens instead of words?

A fixed vocabulary of subword tokens can represent any text, including new words, typos, names, and code, by combining smaller pieces. Using whole words would need an enormous vocabulary and would still fail on words the model has never seen.

Do input and output tokens cost the same?

Not always. Many AI APIs price input tokens and output tokens separately, and output tokens are often more expensive because the model must generate them one at a time.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings