Book 08 · Cheat sheet
AI & Machine Learning
Essential vocabulary for machine learning, large language models and AI-powered applications.
Software Dictionary · softwaredictionary.org/categories/ai/cheat-sheet
- 01AGIArtificial General Intelligence
- AGI (artificial general intelligence) is a hypothetical AI system that could learn and do any intellectual task a person can, not just a narrow set of tasks.
- AGI is a hypothetical AI that can do any intellectual task a person can.
- Today's AI, including LLMs, is narrow AI.
- There is no agreed definition or test for AGI.
- 02AI Agent
- An AI agent is a system that uses an LLM to plan and carry out multi-step tasks by deciding which tools to call, observing the results, and acting again.
- An AI agent uses an LLM to decide on and carry out a sequence of actions.
- It works in a loop: choose an action, call a tool, observe the result, repeat.
- Tools give the agent abilities such as searching, running code, or calling APIs.
- 03AI Alignment
- AI alignment is the field of making AI systems pursue the goals and values their designers intend, so they behave helpfully, honestly, and safely.
- AI alignment aims to make AI systems pursue the goals their designers actually intend.
- Common techniques include fine-tuning, RLHF, principle-based training, and red teaming.
- Reward hacking and sycophancy are examples of misaligned behavior.
- 04Artificial IntelligenceAI
- Artificial intelligence (AI) is the field of computer science that builds systems able to do tasks that normally require human intelligence.
- AI builds systems that do tasks that normally need human intelligence.
- The term dates from 1955, for a workshop at Dartmouth held in 1956.
- Early AI used hand-written rules; modern AI mostly learns from data.
- 05Attention Mechanism
- The attention mechanism is a neural network technique that lets a model decide, for each token, which other parts of the input matter most and focus on them.
- Attention lets each token weigh how relevant every other token is to it.
- It works with queries, keys, and values, combined through dot products and softmax.
- Multi-head attention runs several attention calculations in parallel.
- 06Backpropagation
- Backpropagation is the algorithm that trains neural networks by measuring how much each weight added to the error and nudging every weight to reduce it.
- Backpropagation computes how much each weight contributed to the error.
- It uses the chain rule to pass the error backwards layer by layer.
- One backward pass gives the gradients for all weights at once.
- 07Chain-of-Thought Prompting
- Chain-of-thought prompting is a technique that asks an LLM to reason through intermediate steps before its final answer, improving accuracy on complex tasks.
- Chain-of-thought prompting asks the model to reason step by step before answering.
- Writing intermediate steps gives the model more room to compute and fewer chances to skip logic.
- It helps most with math, logic, planning, and multi-step questions.
- 08Chatbot
- A chatbot is a program that converses with people in text or speech, answering questions or helping with tasks, using scripted rules or a language model.
- A chatbot holds a conversation in text or speech.
- ELIZA, from the mid-1960s, was one of the first chatbots.
- Older chatbots follow scripts; modern ones generate replies with an LLM.
- 09Chunking
- Chunking splits long documents into smaller passages before they are embedded and stored, so a RAG system can find and pass on just the relevant parts.
- Chunking splits documents into passages that are embedded and retrieved one by one.
- Chunks are usually a few hundred tokens long, often with some overlap.
- Splitting at headings, paragraphs or sentences keeps ideas whole.
- 10Computer Vision
- Computer vision is the field of AI that enables computers to interpret images and video, such as recognizing objects, reading text, or detecting faces.
- Computer vision lets software extract meaning from images and video.
- Core tasks are classification, object detection, segmentation, and OCR.
- Images are grids of pixel values processed by neural networks such as CNNs and vision transformers.
- 11Context Engineering
- Context engineering is the practice of choosing what an LLM sees on each call (instructions, documents, tool results, history) so it can do the task reliably.
- Context engineering decides what an LLM sees on each call: instructions, data, tool results and history.
- It is broader than prompt engineering, which focuses on the instruction itself.
- Retrieval, summarization and memory are its main tools in AI agents.
- 12Context Window
- A context window is the maximum amount of text, measured in tokens, that an LLM can consider at once, including the prompt, conversation history, and its reply.
- The context window is the maximum number of tokens an LLM can process in one request.
- It includes the system prompt, history, attached documents, and the model's output.
- LLMs are stateless, so chat apps resend the conversation with each message.
- 13Cosine Similarity
- Cosine similarity measures how alike two vectors are by the angle between them, from -1 to 1; it is the usual way to compare embeddings in semantic search.
- Cosine similarity is the cosine of the angle between two vectors, from -1 to 1.
- It compares direction and ignores length.
- It is the standard way to compare embeddings in semantic search and RAG.
- 14Deep Learning
- Deep learning is a subset of machine learning that uses neural networks with many layers to learn complex patterns from raw data such as images and text.
- Deep learning uses neural networks with many hidden layers.
- Each layer learns a more abstract representation of the data.
- It learns features automatically from raw data such as pixels, audio, or text.
- 15Diffusion Model
- A diffusion model is a generative AI model that makes images, audio, or video by starting from random noise and removing it step by step until content appears.
- A diffusion model generates content by removing noise step by step.
- It is trained by adding noise to real data and learning to predict that noise.
- A text prompt guides the denoising so the output matches the description.
- 16Embedding
- An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- An embedding is a vector of numbers that captures meaning.
- Items with similar meaning have vectors that are close together.
- Cosine similarity is a common way to compare two embeddings.
- 17EvalsEvaluations
- Evals are tests for AI systems: a set of inputs with expected results or grading rules, run after every change to measure how well a model or prompt performs.
- An eval scores an AI system on a fixed set of examples with known good answers or rules.
- Checks range from exact matches and code checks to LLM-as-a-judge and human review.
- They are run after every model, prompt or retrieval change to catch regressions.
- 18Few-Shot Learning
- Few-shot learning is getting an AI model to perform a task from just a handful of examples, most often by placing a few sample inputs and outputs in the prompt.
- Few-shot learning uses a handful of examples to show a model what to do.
- With LLMs, the examples go in the prompt and the model's weights don't change.
- Zero-shot uses no examples, one-shot uses one, and few-shot uses several.
- 19Fine-tuning
- Fine-tuning is the process of taking a pretrained machine learning model and training it further on a smaller, specific dataset to adapt it to one task.
- Fine-tuning continues training a pretrained model on a smaller, task-specific dataset.
- It changes the model's weights, so the new behavior persists without extra prompting.
- Parameter-efficient methods like LoRA train only a small set of extra weights.
- 20Generative AI
- Generative AI is artificial intelligence that creates new content, such as text, images, code, or audio, based on patterns learned from existing data.
- Generative AI creates new text, images, code, audio, or video from a prompt.
- LLMs generate text token by token; many image models use diffusion.
- Outputs are based on learned patterns, not verified facts, so hallucinations happen.
- 21GPTGenerative Pre-trained Transformer
- GPT (Generative Pre-trained Transformer) is OpenAI's family of large language models that generate text by predicting the next token.
- GPT stands for Generative Pre-trained Transformer.
- It is OpenAI's family of large language models, starting with GPT-1 in 2018.
- GPT models are decoder-only transformers that predict the next token.
- 22Gradient Descent
- Gradient descent is an optimization algorithm that trains machine learning models by repeatedly nudging their parameters in the direction that reduces error.
- Gradient descent minimizes a loss function by adjusting a model's parameters step by step.
- Each step moves the parameters opposite to the gradient, the direction in which the error grows fastest.
- The learning rate sets the step size and strongly affects whether training succeeds.
- 23Hallucination
- A hallucination is when an AI model, such as an LLM, confidently produces information that sounds plausible but is false, invented, or unsupported by sources.
- A hallucination is confident but false or made-up AI output.
- It happens because LLMs predict plausible text rather than verify facts.
- Invented citations, APIs, and package names are common examples.
- 24Inference
- Inference is the stage where a trained machine learning model is used to make predictions or generate output from new data, without changing what it learned.
- Inference means using a trained model to make predictions on new data.
- The model's weights do not change during inference.
- Training happens rarely and costs a lot; inference happens on every request.
- 25LLMLarge Language Model
- An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- An LLM generates text by predicting the next token over and over.
- It is trained on massive text datasets and has billions of parameters.
- Its built-in knowledge is frozen at a training cutoff date.
- 26LoRALow-Rank Adaptation
- LoRA is a cheap way to fine-tune a large model: its weights stay frozen and only small added matrices are trained, so a new skill fits in a few megabytes.
- LoRA fine-tunes a model by training small added matrices while the original weights stay frozen.
- It needs far less memory than full fine-tuning and produces a small adapter file.
- Adapters can be swapped on top of one shared base model.
- 27Machine Learning
- Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- Models learn patterns from example data rather than hand-written rules.
- Training adjusts a model's parameters; inference uses the trained model to make predictions.
- The main styles are supervised, unsupervised, and reinforcement learning.
- 28Mixture of ExpertsMoE
- A mixture of experts (MoE) is a neural network design that sends each input to only a few of many small experts, so a huge model costs far less to run.
- MoE layers hold many expert networks and a router that picks a few per token.
- Only the chosen experts run, so compute per token stays small.
- Total parameters can be huge while active parameters stay modest.
- 29Model Context Protocol
- The Model Context Protocol is an open standard that defines how AI applications connect to external tools, data sources, and prompts through a shared interface.
- MCP is an open standard for connecting AI applications to tools and data.
- Servers expose tools, resources, and prompts; host applications connect through clients.
- Messages use JSON-RPC 2.0 over local standard input and output or remote HTTP.
- 30Model Parameters
- Model parameters are the internal numbers, such as weights and biases, that a machine learning model learns in training and uses to turn inputs into outputs.
- Parameters are the learned numbers, mainly weights and biases, inside a model.
- Training adjusts them; inference uses them without changing them.
- Parameter count, such as 7B or 70B, is a rough measure of model size.
- 31Multimodal AI
- Multimodal AI is artificial intelligence that can understand or generate several types of data, such as text, images, audio, and video, in a single model.
- A modality is a type of data, such as text, images, audio, or video.
- Multimodal models map different data types into a shared embedding space.
- They can answer questions about images, read documents, and handle speech.
- 32Natural Language Processing
- Natural language processing is the field of AI that teaches computers to read, understand, and generate human language in the form of text or speech.
- NLP is the area of AI focused on understanding and generating human language.
- Common tasks include translation, sentiment analysis, summarization, and speech recognition.
- Modern NLP relies on tokens, embeddings, and transformer models.
- 33Neural Network
- A neural network is a machine learning model made of layers of connected artificial neurons that learn patterns from data by adjusting numeric weights.
- A neural network is made of layers of connected neurons: input, hidden, and output.
- Each connection has a weight, and learning means adjusting those weights.
- Training uses a loss function and backpropagation to reduce prediction errors.
- 34Overfitting
- Overfitting happens when a machine learning model learns its training data so closely, including its noise, that it performs poorly on new, unseen data.
- An overfit model does well on training data but poorly on new data.
- Always measure performance on held-out data the model never trained on.
- More data, simpler models, regularization, and early stopping reduce overfitting.
- 35Prompt
- A prompt is the input text or instructions you give an AI model, such as an LLM, to tell it what task to perform and what kind of answer you want.
- A prompt is the input that tells an AI model what to do.
- System prompts set the rules; user prompts carry the specific request.
- Clear instructions, context, and examples lead to better answers.
- 36Prompt Engineering
- Prompt engineering is the practice of designing, testing, and refining the instructions given to an AI model so it produces accurate, consistent, useful output.
- Prompt engineering is the process of designing and testing prompts, not just writing one.
- Clear goals, context, constraints, and output formats make results more reliable.
- Few-shot examples and chain-of-thought reasoning are common techniques.
- 37Quantization
- Quantization is a technique that shrinks an AI model by storing its parameters in fewer bits, such as 8 or 4 instead of 16, so inference is faster and cheaper.
- Quantization stores model parameters with fewer bits, such as 8 or 4.
- It cuts memory use sharply and often speeds up inference.
- A scale factor maps the original values to a small range of integers and back.
- 38RAGRetrieval-Augmented Generation
- RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
- RAG retrieves relevant data first, then asks the LLM to answer using it.
- It usually relies on embeddings and a vector index to find relevant text chunks.
- It keeps answers current without retraining the model.
- 39Reasoning Model
- A reasoning model is a language model trained to work through a problem step by step before answering, spending extra computation to do better on hard tasks.
- Reasoning models think step by step before answering.
- They are trained with reinforcement learning on checkable problems.
- OpenAI's o1 in 2024 and DeepSeek-R1 in 2025 popularized them.
- 40Reinforcement Learning
- Reinforcement learning is a type of machine learning in which an agent learns to make decisions by trial and error, earning rewards for good actions.
- An RL agent learns by trial and error from rewards and penalties.
- The core loop is state, action, reward, and next state.
- Agents must balance exploring new actions with exploiting known good ones.
- 41RLHFReinforcement Learning from Human Feedback
- RLHF (reinforcement learning from human feedback) trains a language model to be more helpful and safe using people's judgments of which answers are better.
- RLHF trains a model on human judgments of which answers are better.
- Steps: fine-tune on examples, train a reward model on rankings, then optimize with RL.
- It made assistants such as ChatGPT follow instructions and refuse harmful requests.
- 42Semantic Search
- Semantic search is a search technique that finds results by meaning rather than exact keywords, usually by comparing embeddings of the query and the documents.
- Semantic search matches by meaning, so different wording can still find the right result.
- Queries and documents are compared as embeddings, usually with cosine similarity.
- Always embed queries and documents with the same model.
- 43Supervised Learning
- Supervised learning is machine learning where a model learns from labeled examples, inputs paired with correct answers, to predict outputs for new data.
- Supervised learning trains on labeled data: inputs paired with the correct outputs.
- Classification predicts categories; regression predicts numbers.
- A loss function measures errors, and training adjusts the model to reduce them.
- 44System Prompt
- A system prompt is the instructions an app gives a language model before the conversation starts, setting its role, rules, tone and what it should know.
- The system prompt sets the model's role, rules and context before the chat.
- It is written by the developer and usually hidden from the user.
- Models give it more weight than ordinary user messages.
- 45Temperature
- Temperature is a setting that controls how random an LLM's output is, from focused and predictable at low values to more varied and creative at high values.
- Temperature scales the model's token probabilities before one token is sampled.
- Low values, around 0 to 0.3, give focused and consistent output.
- High values, around 0.8 and above, give more varied, creative, and error-prone output.
- 46Token
- A token is the basic unit of text that an LLM reads and generates, usually a whole word, part of a word, or a punctuation mark, mapped to a numeric ID.
- A token is a chunk of text: a word, part of a word, or a symbol.
- A tokenizer splits text into tokens and maps each one to a numeric ID.
- In English, one token is roughly four characters or three quarters of a word.
- 47Tool Calling
- Tool calling is an LLM feature in which the model asks the application to run a specific function with structured arguments, then uses the result in its answer.
- Tool calling lets a model request a function call with structured arguments.
- Tools are described with a name, a description, and a parameter schema.
- The application, not the model, runs the tool and returns the result.
- 48Training Data
- Training data is the set of examples a machine learning model learns from, and its quality, size, and coverage largely determine how well the model performs.
- Training data is the set of examples a model learns its patterns from.
- Datasets are usually split into training, validation, and test sets.
- Errors, gaps, and biases in the data carry over into the model's behavior.
- 49Transformer
- A transformer is a neural network architecture that uses attention to weigh how each token in a sequence relates to the others, and it powers most modern LLMs.
- A transformer is a neural network architecture built around self-attention.
- Attention lets each token weigh the relevance of every other token in the input.
- Transformers process input tokens in parallel, which makes large-scale training practical.
- 50Unsupervised Learning
- Unsupervised learning is machine learning in which a model finds patterns, groups, or structure in unlabeled data, without being given the correct answers.
- Unsupervised learning finds structure in data that has no labels.
- Clustering, dimensionality reduction, and anomaly detection are the main tasks.
- Results need human interpretation, because there is no correct answer to compare against.
- 51Vector Database
- A vector database is a database designed to store embeddings and quickly find the vectors most similar to a query, which powers semantic search and RAG.
- A vector database stores embeddings and finds the ones most similar to a query.
- It measures similarity with metrics such as cosine similarity or Euclidean distance.
- Approximate nearest neighbor indexes make search fast across millions of vectors.
- 52Vibe Coding
- Vibe coding is building software by describing what you want to an AI and accepting the code it writes, mostly judging the result by whether it seems to work.
- Vibe coding means describing what you want to an AI and accepting its code with little review.
- It is fast for prototypes and lets non-programmers build working software.
- Unread code hides bugs and security holes, which makes it risky for software that must last.