GPU
Graphics Processing Unit
- In Turkish
- ekran kartı
- Pronunciation
- jee-pee-YOO
In short
A GPU (graphics processing unit) is a processor whose thousands of small cores run the same calculation on lots of data at once, for graphics and AI.
What is a GPU?
Drawing a frame means computing the color of millions of pixels, and each pixel follows the same steps. GPUs were designed for exactly that: instead of a few complex cores like a CPU, they have thousands of simpler ones that work in parallel. Games, video editing and 3D rendering depend on them, either as a separate graphics card or integrated into the main processor.
Researchers soon noticed that the same hardware speeds up any math made of many independent operations. NVIDIA's CUDA platform, released in 2007, made GPUs programmable for general computing, and the matrix multiplications at the heart of neural networks turned out to be a perfect fit. Training and running large language models today happens mostly on GPUs and similar AI accelerators.
A GPU has its own fast memory, called VRAM, and a model or dataset usually has to fit in it to run efficiently, which is why memory size is a key figure for AI hardware. Developers rarely program GPUs directly; libraries such as PyTorch, TensorFlow and JAX move the work to the GPU, and graphics programmers use APIs such as Vulkan, Metal and DirectX.
A common misconception is that a GPU makes every program faster. It only helps with work that can be split into many identical, independent pieces. Code full of branches and sequential steps, such as most business logic, runs better on the CPU, and copying data between CPU and GPU memory has its own cost.
Key takeaways
- A GPU has thousands of simple cores that work in parallel.
- It was built for graphics and now also powers AI.
- CUDA, from 2007, made GPUs programmable for general computing.
- VRAM size limits which models and data fit on the GPU.
- Only highly parallel work benefits; branchy code stays on the CPU.
Example
import torch
device = "cuda" if torch.cuda.is_available() else "cpu"
a = torch.rand(8000, 8000, device=device)
b = torch.rand(8000, 8000, device=device)
c = a @ b # one huge matrix multiplication, spread across thousands of GPU cores
print(device, c.shape)Readers ask
Why are GPUs used for AI?
Neural networks are mostly large matrix multiplications, which break into millions of independent operations. A GPU's thousands of cores perform them in parallel, making training and inference far faster than on a CPU.
What is the difference between integrated and dedicated graphics?
Integrated graphics are built into the main processor and share system memory, which saves power. A dedicated graphics card has its own GPU and VRAM and is much faster for games, 3D and AI work.
What is CUDA?
NVIDIA's platform and programming model for running general-purpose computations on its GPUs. Most AI frameworks use CUDA under the hood when running on NVIDIA hardware.
Often compared
See also
- CPUOperating Systems, p. 4A CPU (central processing unit) is the processor that executes a program's instructions, doing the arithmetic, logic and control work all software runs on.
- ParallelismProgramming Fundamentals, p. 42Parallelism is running several computations at literally the same time, on multiple CPU cores, GPUs or machines, so a large job finishes faster.
- Deep LearningAI & Machine Learning, p. 14Deep learning is a subset of machine learning that uses neural networks with many layers to learn complex patterns from raw data such as images and text.
- Neural NetworkAI & Machine Learning, p. 33A neural network is a machine learning model made of layers of connected artificial neurons that learn patterns from data by adjusting numeric weights.
- RAMOperating Systems, p. 25RAM (random access memory) is a computer's fast, temporary working memory, holding the programs and data in use; its contents are lost when power goes off.
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
Spotted a mistake or something missing on this page?Suggest an edit