CPU Cache
- In Turkish
- CPU Önbelleği
- Pronunciation
- see-pee-yoo KASH
In short
A CPU cache is a small, very fast memory on the processor that keeps copies of recently used data from RAM, so the CPU spends less time waiting for memory.
What is a CPU cache?
Main memory is far slower than a modern CPU: fetching a value from RAM can take around 100 nanoseconds, enough time for hundreds of instructions. A CPU cache is a small amount of very fast memory built into the processor that keeps copies of the data and instructions the CPU has used recently or is likely to need next.
Caches are arranged in levels: L1 is the smallest and fastest, tens of kilobytes per core and split into instruction and data caches; L2 is larger and slightly slower; and L3, often tens of megabytes, is shared by all cores. Data moves between memory and cache in fixed-size blocks called cache lines, typically 64 bytes. When the CPU finds what it needs in the cache, it is a cache hit; when it doesn't, a cache miss forces it to fetch the line from the next level or from RAM. Caches work because programs tend to reuse recent data and to access nearby addresses, and multi-core chips use a coherence protocol so every core sees up-to-date values.
Think of a desk in a library. The books you are using sit on your desk (L1), a few more are on a nearby cart (L2), the reading room shelf holds more (L3), and the stacks in the basement are RAM. This is why the CPU cache matters to everyday code: walking through an array in order is much faster than following pointers in a linked list scattered across memory, and two threads that write to different variables on the same cache line can slow each other down, a problem called false sharing.
A CPU cache is often confused with an application cache, such as an in-memory cache in front of a database or an HTTP cache. Both keep copies of data closer to where it is needed, but the CPU cache is managed automatically by the hardware and works on nanosecond timescales, while software caches are managed by your code, which must decide when to invalidate them. CPU registers are different again: they are even smaller and faster storage inside the core itself.
Key takeaways
- A CPU cache is fast memory on the processor that holds copies of data from RAM.
- Caches come in levels: L1 is smallest and fastest, L3 is largest and shared.
- Data moves in cache lines, typically 64 bytes long.
- Sequential, predictable memory access makes the best use of the cache.
- The CPU manages its cache automatically, unlike application-level caches.
Example
#define N 4096
static int grid[N][N];
long sum_rows(void) { /* fast: reads memory in order, */
long s = 0; /* so every 64-byte cache line is fully used */
for (int i = 0; i < N; i++)
for (int j = 0; j < N; j++) s += grid[i][j];
return s;
}
long sum_cols(void) { /* slow: jumps 16 KB on every step, */
long s = 0; /* so almost every read is a cache miss */
for (int j = 0; j < N; j++)
for (int i = 0; i < N; i++) s += grid[i][j];
return s;
}Readers ask
What is the difference between L1, L2, and L3 cache?
They are levels of cache that trade size for speed. L1 is the smallest and fastest and belongs to one core, L2 is larger and a little slower, and L3 is the largest and slowest cache level, usually shared by all cores.
What is a cache miss?
A cache miss happens when the data the CPU needs is not in the cache, so it has to be fetched from a slower cache level or from RAM. Frequent misses can make a program several times slower.
Does CPU cache size matter?
It matters for workloads whose active data almost fits in the cache, such as games, databases, and scientific code. A larger cache means fewer trips to RAM, but access patterns in the code often matter more than raw size.
See also
- CacheBackend & APIs, p. 8A cache is a fast, temporary storage layer that keeps copies of frequently used data so later requests can be served quickly without repeating slow work.
- ArrayProgramming Fundamentals, p. 3An array is an ordered collection of values stored under one name, where each item is accessed by its numeric position, called an index, usually starting at 0.
- Linked ListData Structures, p. 22A linked list is a data structure that stores items in separate nodes, where each node holds a value and a reference to the next node in the chain.
- Context SwitchOperating Systems, p. 3A context switch is when the operating system saves the state of the running thread or process and restores another one's state so it can use the CPU.
- PagingOperating Systems, p. 22Paging is a memory management scheme that splits memory into fixed-size pages and uses page tables to map each process's virtual pages to physical RAM.
- Big O NotationProgramming Fundamentals, p. 6Big O notation describes how an algorithm's running time or memory use grows as its input gets larger, focusing on the growth rate rather than exact speed.
- CPUOperating Systems, p. 4A CPU (central processing unit) is the processor that executes a program's instructions, doing the arithmetic, logic and control work all software runs on.
- RAMOperating Systems, p. 25RAM (random access memory) is a computer's fast, temporary working memory, holding the programs and data in use; its contents are lost when power goes off.
Spotted a mistake or something missing on this page?Suggest an edit