Storage pyramid -- why so many layers
Have you ever wondered why computers have both "memory" and "hard drive"? Why not put all data in the fastest place?
The answer lies in the design of computer storage systems--Storage pyramidThis lecture gives you a thorough understanding of the architect's most core trade-off wisdom.
Everyday analogy: kitchen vs. supermarket warehouse
Use the placement of ingredients while cooking to build an intuitive impression first.
Imagine you are cooking.
Salt and soy sauce, you put them by the stove, within reach--these are the computer'sregisterandL1 cache。
Ingredients in the fridge, you need to walk a few steps to get them--this is equivalent toMemory (RAM)。
Rice, flour, and oil stored in the basement, you have to make a special trip--this isSolid State Drive (SSD)。
And the stock in the supermarket warehouse, you wouldn't keep it at home at all--that isHard Disk Drive (HDD)and cloud storage.
This reveals a simple rule:Things used more often are placed closer, but closer space is limited and more expensive; things used less often are placed farther away, with larger capacity and also cheaper.
Computer storage systems are designed according to this simple rule.
Storage pyramid: the trade-off between speed and capacity
Inside a computer, storage devices form a pyramid structure according to the progressive relationship of "speed-capacity-cost".
example Storage pyramidallocateGraph — diagram-design Specification SVG:纸色底 + emit丝LineDividelayer,锈red accent onlymarkNote顶layerregisterStorage pyramid diagram: the higher up, the faster and smaller; the lower down, the slower and larger. Hover over a trapezoid to see details for each layer.
Core rules
This pyramid diagram reveals an iron rule in hardware design:
- Faster, more expensive, smallerRegisters are inside the CPU, made of the fastest transistors, but only a few kilobytes in size. Because chip area is extremely expensive.
- Slower, cheaper, largerHard drives use magnetic heads and spinning platters, with extremely low cost, and can hold tens of terabytes of data.
- Purpose of the pyramid structureUsing a combination of "small and fast + large and slow" achieves an experience close to the fastest storage at a reasonable cost.
If a computer used registers for all storage, not only would the price be astronomical, the chip area would also be too large to manufacture.
Conversely, if it used only hard drives, the computer would be too slow to use--you would have to wait hundreds of milliseconds to open any program.
Detailed comparison of each layer
Put the key attributes of the seven storage layers into one table for easy horizontal comparison.
| Tier | Typical capacity | Access latency | Relative speed | Physical location | Manufacturing cost |
|---|---|---|---|---|---|
| register | ~1 KB | ~0.3 ns | 1x (baseline) | Inside the CPU core | extremely high |
| L1 cache | ~64 KB | ~1 ns | 3x slower than registers | Inside the CPU core | extremely high |
| L2 cache | ~256 KB | ~4 ns | 13x slower | Inside the CPU core | Very High |
| L3 cache | ~8 MB | ~12 ns | 40x slower | On the CPU chip, shared by multiple cores | High |
| Memory (RAM) | ~16 GB | ~100 ns | 333x slower | On the motherboard (separate chip) | Medium |
| Solid State Drive (SSD) | ~1 TB | ~100 μs | 333,333x slower | Inside the computer case (separate device) | Relatively Low |
| Hard Disk Drive (HDD) | ~10 TB | ~10 ms | 33,333,333x slower | Inside the computer case (separate device) | Low |
Notice the order-of-magnitude jumps: memory is 100 times slower than L1 cache, SSD is 1000 times slower than memory, and HDD is 100 times slower than SSD. The speed gap between every layer is measured in orders of magnitude.
Why is SSD so much faster than HDD?
The key difference lies in whether there are mechanical parts.
SSD uses flash memory chips, reading purely electronically, with no mechanical parts.
HDD uses spinning platters and a moving head--to read a piece of data, the head must first move to the correct track (seek time, about 5-10 ms), then wait for the platter to rotate to the correct position (rotational latency, about 2-4 ms).
This mechanical movement process is exactly the root cause of HDD being hundreds of times slower than SSD.
Interactive demo: simulating access times of multi-layer storage
Use a Python snippet to simulate the latency differences when reading the same address from each of the seven storage layers.
Example
Memory Pyramid Access Latency Simulator (example Demo)
Simulate reading data blocks of the same size from different levels to intuitively compare the speed differences of each level
"""
class StorageHierarchy:
"""Simulate the memory hierarchy of a computer"""
def __init__(self):
# Simulated latency (in nanoseconds) and typical capacity of each storage layer
self.layers = {
'Register': {'latency_ns': 0.3, 'capacity': '~1 KB', 'color': 'red'},
'L1 Cache': {'latency_ns': 1, 'capacity': '~64 KB', 'color': 'orange'},
'L2 Cache': {'latency_ns': 4, 'capacity': '~256 KB', 'color': 'gold'},
'L3 Cache': {'latency_ns': 12, 'capacity': '~8 MB', 'color': 'yellow'},
'RAM': {'latency_ns': 100, 'capacity': '~16 GB', 'color': 'green'},
'SSD': {'latency_ns': 100000, 'capacity': '~1 TB', 'color': 'blue'},
'HDD': {'latency_ns': 10000000, 'capacity': '~10 TB', 'color': 'purple'},
}
def read(self, layer_name, address):
"""
Simulate reading data from the specified layer
Return simulated data values
"""
value = f"DATA_FROM_{layer_name.upper()}_{address}"
return value
def compare_access(self, address, num_accesses=3):
"""Compare the speed of reading the same address from each level."""
print("=" * 65)
print(f"Memory hierarchy access latency comparison (address: {address})")
print("=" * 65)
print(f{'Level':<12} {'Latency (ns)':>12} {'Capacity':<12} {'Relative to register':>12})
print("-" * 65)
baseline = self.layers['Register']['latency_ns']
for name, info in self.layers.items():
lat_ns = info['latency_ns']
ratio = lat_ns / baseline
cap = info['capacity']
print(f"{name:<12} {lat_ns:>10,.1f} ns {cap:<12} {ratio:>10,.0f}x")
# Run demo
print(EXAMPLE Storage System Tutorial: Storage Pyramid Access Latency Comparison)
print()
store = StorageHierarchy()
store.compare_access("0x7FFF1234")
print()
print("=" * 65)
print("Conclusion analysis:")
print("=" * 65)
# Calculate key ratios
register_lat = 0.3 # ns
ram_lat = 100 # ns
ssd_lat = 100000 # ns
hdd_lat = 10000000 # ns
print(f1. Memory (RAM) is {ram_lat / register_lat:,.0f} times slower than registers)
print(f2. SSD is {ssd_lat / ram_lat:,.0f} times slower than memory (RAM))
print(f3. HDD is {hdd_lat / ram_lat:,.0f} times slower than memory (RAM))
print(f4. HDD is {hdd_lat / register_lat:,.0f} times slower than registers)
print()
print("If register access takes 1 second, then:")
print(f- Fetching from L1 cache takes {1/register_lat:.0f} seconds)
print(f" - fromMemoryGetrequires {ram_lat/register_lat:,.0f} second(约 {ram_lat/register_lat/60:.0f} Divide钟)")
print(f" - from HDD Getrequires {hdd_lat/register_lat:,.0f} second(约 {hdd_lat/register_lat/3600:,.0f} hours)")
Run result:
EXAMPLE 存储系统教学: 存储金字塔访问延迟对比 ================================================================= 存储层次访问延迟对比(地址: 0x7FFF1234) ================================================================= 层级 延迟(纳秒) 容量 相对寄存器 ----------------------------------------------------------------- Register 0.3 ns ~1 KB 1x L1 Cache 1.0 ns ~64 KB 3x L2 Cache 4.0 ns ~256 KB 13x L3 Cache 12.0 ns ~8 MB 40x RAM 100.0 ns ~16 GB 333x SSD 100,000.0 ns ~1 TB 333,333x HDD 10,000,000.0 ns ~10 TB 33,333,333x ================================================================= 结论分析: ================================================================= 1. 内存(RAM) 比 寄存器 慢 333 倍 2. SSD 比 内存(RAM) 慢 1,000 倍 3. HDD 比 内存(RAM) 慢 100,000 倍 4. HDD 比 寄存器 慢 33,333,333 倍 如果寄存器访问数据需要 1 秒,那么: - 从 L1 缓存获取需要 3 秒 - 从内存获取需要 333 秒(约 6 分钟) - 从 HDD 获取需要 33,333,333 秒(约 9,259 小时)
Interactive demo: logarithmic comparison chart of access latency across layers
The chart below draws a bar chart on a logarithmic scale; adjacent intervals on the horizontal axis differ by 10 times. Hover over a bar to see capacity and latency details.
example storage hierarchy latency comparison — Plotly log-coordinate horizontal bar chart (standard selection: use Plotly for logarithmic coordinate system)How the pyramid works: layer-by-layer caching strategy
The pyramid is not static--it has an automatic data movement mechanism.
Data flow rules
- When the CPU Needs DataFirst look in the fastest L1 cache; if not found, go to L2; if still not found, go to L3, and so on down to main memory.
- After reading data from the slow layerIt not only gives the data to the CPU, but also stores a copy in a faster layer. That way the next use is fast.
- What to do when the fast layer is full?According to a certain policy (such as LRU--evicting the least recently used), kick infrequently used data back to slower layers to make room for new data.
This process is completely transparent to programmers--you don't need to manually manage which cache layer when writing programs; hardware and the operating system automatically do all of this.
Real-world examples
Suppose you are editing a video file:
- Video files existHDD or SSDAbove (bottom layer).
- When you open a file, the operating system loads part of it intoMemory (RAM)In.
- When you start playing, the CPU copies the data of the frames currently being processed toL3/L2/L1 cacheIn.
- Pixel values currently being computed by the ALU are stored inregisterIn.
When you feel "very smooth" in editing software, it's because the CPU finds the data it needs in the cache most of the time.
Historical background: why the pyramid has this shape
The number of pyramid layers is a historical product of the ever-widening speed gap between CPU and memory.
In the 1980s, the speed gap between CPU and memory was not large. But as semiconductor manufacturing advanced, CPU speed grew at about 60% per year, while memory speed only grew at about 10% per year.
This ever-widening gap is called「Memory Wall」(Memory Wall)。
The storage pyramid is engineers' response to the "memory wall"--using multiple levels of cache to buffer the speed gap between CPU and main memory.
| era | CPU frequency | Memory latency | Speed gap | Cache hierarchy |
|---|---|---|---|---|
| 1980s | ~10 MHz | ~200 ns | approximately 2x | None or Level 1 |
| 1990s | ~200 MHz | ~70 ns | About 14 times | L1 + L2 |
| 2000s | ~3 GHz | ~50 ns | About 150 times | L1 + L2 + L3 |
| 2020s | ~5 GHz | ~80 ns | About 400 times | L1 + L2 + L3 |
other extensionsNote that memory latency has barely changed in essence over 40 years! It's not that memory hasn't improved, but that CPU has improved too fast. The physical limit of memory (capacitor charge/discharge speed) makes it very difficult to significantly shorten its latency.