Skip to main content

Book 08 · Cheat sheet

AI & Machine Learning

Essential vocabulary for machine learning, large language models and AI-powered applications.

01AGIArtificial General Intelligence
AGI (artificial general intelligence) is a hypothetical AI system that could learn and do any intellectual task a person can, not just a narrow set of tasks.
  • AGI is a hypothetical AI that can do any intellectual task a person can.
  • Today's AI, including LLMs, is narrow AI.
  • There is no agreed definition or test for AGI.
02AI Agent
An AI agent is a system that uses an LLM to plan and carry out multi-step tasks by deciding which tools to call, observing the results, and acting again.
  • An AI agent uses an LLM to decide on and carry out a sequence of actions.
  • It works in a loop: choose an action, call a tool, observe the result, repeat.
  • Tools give the agent abilities such as searching, running code, or calling APIs.
03AI Alignment
AI alignment is the field of making AI systems pursue the goals and values their designers intend, so they behave helpfully, honestly, and safely.
  • AI alignment aims to make AI systems pursue the goals their designers actually intend.
  • Common techniques include fine-tuning, RLHF, principle-based training, and red teaming.
  • Reward hacking and sycophancy are examples of misaligned behavior.
04Artificial IntelligenceAI
Artificial intelligence (AI) is the field of computer science that builds systems able to do tasks that normally require human intelligence.
  • AI builds systems that do tasks that normally need human intelligence.
  • The term dates from 1955, for a workshop at Dartmouth held in 1956.
  • Early AI used hand-written rules; modern AI mostly learns from data.
05Attention Mechanism
The attention mechanism is a neural network technique that lets a model decide, for each token, which other parts of the input matter most and focus on them.
  • Attention lets each token weigh how relevant every other token is to it.
  • It works with queries, keys, and values, combined through dot products and softmax.
  • Multi-head attention runs several attention calculations in parallel.
06Backpropagation
Backpropagation is the algorithm that trains neural networks by measuring how much each weight added to the error and nudging every weight to reduce it.
  • Backpropagation computes how much each weight contributed to the error.
  • It uses the chain rule to pass the error backwards layer by layer.
  • One backward pass gives the gradients for all weights at once.
07Chain-of-Thought Prompting
Chain-of-thought prompting is a technique that asks an LLM to reason through intermediate steps before its final answer, improving accuracy on complex tasks.
  • Chain-of-thought prompting asks the model to reason step by step before answering.
  • Writing intermediate steps gives the model more room to compute and fewer chances to skip logic.
  • It helps most with math, logic, planning, and multi-step questions.
08Chatbot
A chatbot is a program that converses with people in text or speech, answering questions or helping with tasks, using scripted rules or a language model.
  • A chatbot holds a conversation in text or speech.
  • ELIZA, from the mid-1960s, was one of the first chatbots.
  • Older chatbots follow scripts; modern ones generate replies with an LLM.
09Chunking
Chunking splits long documents into smaller passages before they are embedded and stored, so a RAG system can find and pass on just the relevant parts.
  • Chunking splits documents into passages that are embedded and retrieved one by one.
  • Chunks are usually a few hundred tokens long, often with some overlap.
  • Splitting at headings, paragraphs or sentences keeps ideas whole.
10Computer Vision
Computer vision is the field of AI that enables computers to interpret images and video, such as recognizing objects, reading text, or detecting faces.
  • Computer vision lets software extract meaning from images and video.
  • Core tasks are classification, object detection, segmentation, and OCR.
  • Images are grids of pixel values processed by neural networks such as CNNs and vision transformers.
11Context Engineering
Context engineering is the practice of choosing what an LLM sees on each call (instructions, documents, tool results, history) so it can do the task reliably.
  • Context engineering decides what an LLM sees on each call: instructions, data, tool results and history.
  • It is broader than prompt engineering, which focuses on the instruction itself.
  • Retrieval, summarization and memory are its main tools in AI agents.
12Context Window
A context window is the maximum amount of text, measured in tokens, that an LLM can consider at once, including the prompt, conversation history, and its reply.
  • The context window is the maximum number of tokens an LLM can process in one request.
  • It includes the system prompt, history, attached documents, and the model's output.
  • LLMs are stateless, so chat apps resend the conversation with each message.
13Cosine Similarity
Cosine similarity measures how alike two vectors are by the angle between them, from -1 to 1; it is the usual way to compare embeddings in semantic search.
  • Cosine similarity is the cosine of the angle between two vectors, from -1 to 1.
  • It compares direction and ignores length.
  • It is the standard way to compare embeddings in semantic search and RAG.
14Deep Learning
Deep learning is a subset of machine learning that uses neural networks with many layers to learn complex patterns from raw data such as images and text.
  • Deep learning uses neural networks with many hidden layers.
  • Each layer learns a more abstract representation of the data.
  • It learns features automatically from raw data such as pixels, audio, or text.
15Diffusion Model
A diffusion model is a generative AI model that makes images, audio, or video by starting from random noise and removing it step by step until content appears.
  • A diffusion model generates content by removing noise step by step.
  • It is trained by adding noise to real data and learning to predict that noise.
  • A text prompt guides the denoising so the output matches the description.
16Embedding
An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
  • An embedding is a vector of numbers that captures meaning.
  • Items with similar meaning have vectors that are close together.
  • Cosine similarity is a common way to compare two embeddings.
17EvalsEvaluations
Evals are tests for AI systems: a set of inputs with expected results or grading rules, run after every change to measure how well a model or prompt performs.
  • An eval scores an AI system on a fixed set of examples with known good answers or rules.
  • Checks range from exact matches and code checks to LLM-as-a-judge and human review.
  • They are run after every model, prompt or retrieval change to catch regressions.
18Few-Shot Learning
Few-shot learning is getting an AI model to perform a task from just a handful of examples, most often by placing a few sample inputs and outputs in the prompt.
  • Few-shot learning uses a handful of examples to show a model what to do.
  • With LLMs, the examples go in the prompt and the model's weights don't change.
  • Zero-shot uses no examples, one-shot uses one, and few-shot uses several.
19Fine-tuning
Fine-tuning is the process of taking a pretrained machine learning model and training it further on a smaller, specific dataset to adapt it to one task.
  • Fine-tuning continues training a pretrained model on a smaller, task-specific dataset.
  • It changes the model's weights, so the new behavior persists without extra prompting.
  • Parameter-efficient methods like LoRA train only a small set of extra weights.
20Generative AI
Generative AI is artificial intelligence that creates new content, such as text, images, code, or audio, based on patterns learned from existing data.
  • Generative AI creates new text, images, code, audio, or video from a prompt.
  • LLMs generate text token by token; many image models use diffusion.
  • Outputs are based on learned patterns, not verified facts, so hallucinations happen.
21GPTGenerative Pre-trained Transformer
GPT (Generative Pre-trained Transformer) is OpenAI's family of large language models that generate text by predicting the next token.
  • GPT stands for Generative Pre-trained Transformer.
  • It is OpenAI's family of large language models, starting with GPT-1 in 2018.
  • GPT models are decoder-only transformers that predict the next token.
22Gradient Descent
Gradient descent is an optimization algorithm that trains machine learning models by repeatedly nudging their parameters in the direction that reduces error.
  • Gradient descent minimizes a loss function by adjusting a model's parameters step by step.
  • Each step moves the parameters opposite to the gradient, the direction in which the error grows fastest.
  • The learning rate sets the step size and strongly affects whether training succeeds.
23Hallucination
A hallucination is when an AI model, such as an LLM, confidently produces information that sounds plausible but is false, invented, or unsupported by sources.
  • A hallucination is confident but false or made-up AI output.
  • It happens because LLMs predict plausible text rather than verify facts.
  • Invented citations, APIs, and package names are common examples.
24Inference
Inference is the stage where a trained machine learning model is used to make predictions or generate output from new data, without changing what it learned.
  • Inference means using a trained model to make predictions on new data.
  • The model's weights do not change during inference.
  • Training happens rarely and costs a lot; inference happens on every request.
25LLMLarge Language Model
An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
  • An LLM generates text by predicting the next token over and over.
  • It is trained on massive text datasets and has billions of parameters.
  • Its built-in knowledge is frozen at a training cutoff date.
26LoRALow-Rank Adaptation
LoRA is a cheap way to fine-tune a large model: its weights stay frozen and only small added matrices are trained, so a new skill fits in a few megabytes.
  • LoRA fine-tunes a model by training small added matrices while the original weights stay frozen.
  • It needs far less memory than full fine-tuning and produces a small adapter file.
  • Adapters can be swapped on top of one shared base model.
27Machine Learning
Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
  • Models learn patterns from example data rather than hand-written rules.
  • Training adjusts a model's parameters; inference uses the trained model to make predictions.
  • The main styles are supervised, unsupervised, and reinforcement learning.
28Mixture of ExpertsMoE
A mixture of experts (MoE) is a neural network design that sends each input to only a few of many small experts, so a huge model costs far less to run.
  • MoE layers hold many expert networks and a router that picks a few per token.
  • Only the chosen experts run, so compute per token stays small.
  • Total parameters can be huge while active parameters stay modest.
29Model Context Protocol
The Model Context Protocol is an open standard that defines how AI applications connect to external tools, data sources, and prompts through a shared interface.
  • MCP is an open standard for connecting AI applications to tools and data.
  • Servers expose tools, resources, and prompts; host applications connect through clients.
  • Messages use JSON-RPC 2.0 over local standard input and output or remote HTTP.
30Model Parameters
Model parameters are the internal numbers, such as weights and biases, that a machine learning model learns in training and uses to turn inputs into outputs.
  • Parameters are the learned numbers, mainly weights and biases, inside a model.
  • Training adjusts them; inference uses them without changing them.
  • Parameter count, such as 7B or 70B, is a rough measure of model size.
31Multimodal AI
Multimodal AI is artificial intelligence that can understand or generate several types of data, such as text, images, audio, and video, in a single model.
  • A modality is a type of data, such as text, images, audio, or video.
  • Multimodal models map different data types into a shared embedding space.
  • They can answer questions about images, read documents, and handle speech.
32Natural Language Processing
Natural language processing is the field of AI that teaches computers to read, understand, and generate human language in the form of text or speech.
  • NLP is the area of AI focused on understanding and generating human language.
  • Common tasks include translation, sentiment analysis, summarization, and speech recognition.
  • Modern NLP relies on tokens, embeddings, and transformer models.
33Neural Network
A neural network is a machine learning model made of layers of connected artificial neurons that learn patterns from data by adjusting numeric weights.
  • A neural network is made of layers of connected neurons: input, hidden, and output.
  • Each connection has a weight, and learning means adjusting those weights.
  • Training uses a loss function and backpropagation to reduce prediction errors.
34Overfitting
Overfitting happens when a machine learning model learns its training data so closely, including its noise, that it performs poorly on new, unseen data.
  • An overfit model does well on training data but poorly on new data.
  • Always measure performance on held-out data the model never trained on.
  • More data, simpler models, regularization, and early stopping reduce overfitting.
35Prompt
A prompt is the input text or instructions you give an AI model, such as an LLM, to tell it what task to perform and what kind of answer you want.
  • A prompt is the input that tells an AI model what to do.
  • System prompts set the rules; user prompts carry the specific request.
  • Clear instructions, context, and examples lead to better answers.
36Prompt Engineering
Prompt engineering is the practice of designing, testing, and refining the instructions given to an AI model so it produces accurate, consistent, useful output.
  • Prompt engineering is the process of designing and testing prompts, not just writing one.
  • Clear goals, context, constraints, and output formats make results more reliable.
  • Few-shot examples and chain-of-thought reasoning are common techniques.
37Quantization
Quantization is a technique that shrinks an AI model by storing its parameters in fewer bits, such as 8 or 4 instead of 16, so inference is faster and cheaper.
  • Quantization stores model parameters with fewer bits, such as 8 or 4.
  • It cuts memory use sharply and often speeds up inference.
  • A scale factor maps the original values to a small range of integers and back.
38RAGRetrieval-Augmented Generation
RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
  • RAG retrieves relevant data first, then asks the LLM to answer using it.
  • It usually relies on embeddings and a vector index to find relevant text chunks.
  • It keeps answers current without retraining the model.
39Reasoning Model
A reasoning model is a language model trained to work through a problem step by step before answering, spending extra computation to do better on hard tasks.
  • Reasoning models think step by step before answering.
  • They are trained with reinforcement learning on checkable problems.
  • OpenAI's o1 in 2024 and DeepSeek-R1 in 2025 popularized them.
40Reinforcement Learning
Reinforcement learning is a type of machine learning in which an agent learns to make decisions by trial and error, earning rewards for good actions.
  • An RL agent learns by trial and error from rewards and penalties.
  • The core loop is state, action, reward, and next state.
  • Agents must balance exploring new actions with exploiting known good ones.
41RLHFReinforcement Learning from Human Feedback
RLHF (reinforcement learning from human feedback) trains a language model to be more helpful and safe using people's judgments of which answers are better.
  • RLHF trains a model on human judgments of which answers are better.
  • Steps: fine-tune on examples, train a reward model on rankings, then optimize with RL.
  • It made assistants such as ChatGPT follow instructions and refuse harmful requests.
42Semantic Search
Semantic search is a search technique that finds results by meaning rather than exact keywords, usually by comparing embeddings of the query and the documents.
  • Semantic search matches by meaning, so different wording can still find the right result.
  • Queries and documents are compared as embeddings, usually with cosine similarity.
  • Always embed queries and documents with the same model.
43Supervised Learning
Supervised learning is machine learning where a model learns from labeled examples, inputs paired with correct answers, to predict outputs for new data.
  • Supervised learning trains on labeled data: inputs paired with the correct outputs.
  • Classification predicts categories; regression predicts numbers.
  • A loss function measures errors, and training adjusts the model to reduce them.
44System Prompt
A system prompt is the instructions an app gives a language model before the conversation starts, setting its role, rules, tone and what it should know.
  • The system prompt sets the model's role, rules and context before the chat.
  • It is written by the developer and usually hidden from the user.
  • Models give it more weight than ordinary user messages.
45Temperature
Temperature is a setting that controls how random an LLM's output is, from focused and predictable at low values to more varied and creative at high values.
  • Temperature scales the model's token probabilities before one token is sampled.
  • Low values, around 0 to 0.3, give focused and consistent output.
  • High values, around 0.8 and above, give more varied, creative, and error-prone output.
46Token
A token is the basic unit of text that an LLM reads and generates, usually a whole word, part of a word, or a punctuation mark, mapped to a numeric ID.
  • A token is a chunk of text: a word, part of a word, or a symbol.
  • A tokenizer splits text into tokens and maps each one to a numeric ID.
  • In English, one token is roughly four characters or three quarters of a word.
47Tool Calling
Tool calling is an LLM feature in which the model asks the application to run a specific function with structured arguments, then uses the result in its answer.
  • Tool calling lets a model request a function call with structured arguments.
  • Tools are described with a name, a description, and a parameter schema.
  • The application, not the model, runs the tool and returns the result.
48Training Data
Training data is the set of examples a machine learning model learns from, and its quality, size, and coverage largely determine how well the model performs.
  • Training data is the set of examples a model learns its patterns from.
  • Datasets are usually split into training, validation, and test sets.
  • Errors, gaps, and biases in the data carry over into the model's behavior.
49Transformer
A transformer is a neural network architecture that uses attention to weigh how each token in a sequence relates to the others, and it powers most modern LLMs.
  • A transformer is a neural network architecture built around self-attention.
  • Attention lets each token weigh the relevance of every other token in the input.
  • Transformers process input tokens in parallel, which makes large-scale training practical.
50Unsupervised Learning
Unsupervised learning is machine learning in which a model finds patterns, groups, or structure in unlabeled data, without being given the correct answers.
  • Unsupervised learning finds structure in data that has no labels.
  • Clustering, dimensionality reduction, and anomaly detection are the main tasks.
  • Results need human interpretation, because there is no correct answer to compare against.
51Vector Database
A vector database is a database designed to store embeddings and quickly find the vectors most similar to a query, which powers semantic search and RAG.
  • A vector database stores embeddings and finds the ones most similar to a query.
  • It measures similarity with metrics such as cosine similarity or Euclidean distance.
  • Approximate nearest neighbor indexes make search fast across millions of vectors.
52Vibe Coding
Vibe coding is building software by describing what you want to an AI and accepting the code it writes, mostly judging the result by whether it seems to work.
  • Vibe coding means describing what you want to an AI and accepting its code with little review.
  • It is fast for prototypes and lets non-programmers build working software.
  • Unread code hides bugs and security holes, which makes it risky for software that must last.
52 terms from Software Dictionary. Full explanations, examples and FAQs at softwaredictionary.org/categories/ai

Back to the bookTip: pick "Save as PDF" in the print dialog to keep a copy.

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings