Shelf 1 · Computing Foundations · 3 / 45
Why AI Needs GPUs and Memory: The Roles of CPUs, GPUs, and HBM
How CPU latency, GPU throughput, RAM, VRAM, HBM, training state, inference, and the KV cache shape AI computers, and where computing goes next.
Check this article’s sources (11)Article brief
Behind a single short reply to a prompt, CPUs, GPUs, memory, and networks each do a different job while moving enormous amounts of data.
A useful mental model
Picture the CPU as a stage manager making control decisions, the GPU as a large ensemble repeating structured work in parallel, memory as the worktables holding scores and intermediate results, and interconnects as the transport routes.
Where the analogy stops
This analogy explains roles, not a universal ranking. Real performance changes with the model, numeric precision, batch size, software, memory capacity, bandwidth and latency, communication, and power. Adding GPUs or memory does not always make a workload faster.
You will be able to say what CPUs, GPUs, and memory are each good at without ranking them on one speed table, and locate an AI bottleneck across the whole system.
Open the glossaryArticle contents13 chaptersJump to a chapter
1The 30-second answer: CPUs, GPUs, and memory play different roles
| Element | Strength | Typical AI-system role |
|---|---|---|
| CPU | Low latency on a few threads, complex branching, broad compatibility | OS, I/O, preprocessing, scheduling, accelerator control |
| GPU / AI accelerator | High throughput across many parallel operations | Matrix and tensor work, training, batched inference |
| Memory | A hierarchy of capacity, bandwidth, and latency | Weights, activations, gradients, optimizer state, and KV cache |
| Interconnect | Moves data between devices | Synchronizing and distributing work across accelerators |
A GPU is not always faster than a CPU, and more memory does not make a model smarter. Any comparison has to weigh the workload, the software, the numerical precision, and the data movement together.
2The shape of the work matters more than a performance ranking
Latency is how long one job takes to finish; throughput is how much work completes per unit of time. A CPU is built for a fast response along one or a few complex paths. A GPU is built for the combined rate of many similar ones.
Small inputs, heavy branching, sequential dependencies, or transfers between host and device can leave GPU lanes idle. When the same large matrices are processed over and over, many GPU lanes together can outrun a handful of powerful CPU cores. The first question is not which processor is “better,” but how regularly the work divides.
3CPU strengths: latency, branching, and control
A modern high-performance CPU core combines branch prediction, out-of-order execution, large caches, and a broad instruction set to handle shifting control flow with low latency. That suits responsive work: the operating system, networking, storage, databases, compression, tokenization, data loaders, and error handling.
CPUs also have SIMD units, matrix extensions, and multiple cores, so they can run AI inference themselves. That is a reasonable choice when the model is small, the batch is tiny, the latency target is strict, or moving data to an accelerator costs more than it saves. Even in accelerator-heavy systems, the CPU still prepares jobs, launches kernels, manages I/O, and handles failures.
4GPU strengths: parallelism and throughput
A GPU spreads one kernel across many lightweight threads to sustain high throughput on regular arithmetic. The model grew out of pixel and vertex work, and it maps well onto matrix and tensor operations over wide data. Dedicated matrix units can drop to lower precision to run the multiply-accumulate patterns that dominate AI.
Branch divergence, tiny jobs, frequent round trips to the CPU, or memory that cannot keep up will leave those lanes waiting. Hardware parallelism is usable only when compilers, libraries, and toolchains such as CUDA support it. Peak FLOPS or TOPS alone cannot tell you how an application will perform.
5Why AI uses so much matrix math
Many neural-network layers multiply an input vector by a weight matrix and then apply a nonlinear transformation. In Transformers, both attention and the feed-forward networks contain large matrix multiplications. Because the same multiply-accumulate structure repeats, the work fits the parallel units in GPUs and TPUs.
AI is not just matrix multiplication. Normalization, activation, sampling, routing, preprocessing, and communication all matter. Training behaves differently from inference, and prompt prefill behaves differently from token-by-token decode. A benchmark is meaningful only when the model, the precision, the batch size, the quality target, and the scenario line up.
6Memory: keep capacity, bandwidth, and latency separate
In conventional architectures, memory does not do the arithmetic. It stores data and delivers it to the units that do. Capacity decides what fits at once, bandwidth decides how many bytes move per second, and latency decides how long a single request waits. The three are independent: a large pool can be slow, and wide bandwidth does not remove random-access latency.
| Layer | Location and role |
|---|---|
| Registers / on-chip SRAM / cache | Small, close to arithmetic, and fast |
| VRAM / HBM | Large accelerator-local memory; HBM uses stacking and a wide interface for bandwidth |
| System DRAM | CPU-side main memory with capacity and generality |
| SSD / storage | Larger and slower than execution memory; used for loading and offload |
7Training must store more than weights
Training holds gradients, optimizer state, and forward-pass activations on top of the model parameters. How much that adds depends on the precision and the optimizer, so multiplying the parameter count by a byte size understates it. ZeRO partitions optimizer state, gradients, and parameters across devices for exactly this reason: the memory bill constrains the design of the whole system.
Mixed precision computes and stores some values at lower precision, which improves throughput and effective capacity but calls for techniques that keep the numerics stable. Activation checkpointing recomputes selected activations instead of keeping them, trading compute for memory. Model parallelism raises the total capacity available but adds communication and synchronization.
8Inference holds weights and a growing KV cache
LLM inference reads the model weights again and again as it generates tokens. Autoregressive decode may do little new arithmetic per step, so it runs up against bandwidth rather than compute. Prompt prefill and larger batches can have higher arithmetic intensity, so “all AI inference is memory-bound” is too broad.
The KV cache keeps attention keys and values rather than recomputing them, and it grows with the number of layers, tokens, and concurrent requests. PagedAttention proposed managing that cache the way an operating system manages virtual memory, cutting the waste in contiguous allocation. Long context is therefore a question of memory capacity and bandwidth, not only of algorithmic capability.
9How many gigabytes does it take just to hold 7B or 70B weights?
Counting weights alone at two bytes per FP16 or BF16 parameter gives about 14 GB for 7 billion parameters and 140 GB for 70 billion. A theoretical packed 4-bit weight body would come to roughly 3.5 GB and 35 GB. These are decimal gigabytes.
Real systems then add quantization metadata, values kept at higher precision, runtime and allocator headroom, temporary buffers, and the KV cache. Training adds gradients, optimizer state, and activations on top. The arithmetic does not tell you how many GPUs to buy; it is a floor that shows why capacity decides where a model can run.
10Fast arithmetic still waits when the data cannot arrive
The Roofline model bounds performance by the smaller of two quantities: peak compute, and memory bandwidth multiplied by operational intensity, the operations performed per byte moved. A kernel that moves a lot of data without much arithmetic hits the bandwidth roof, and adding arithmetic units does not help.
FlashAttention returns the same attention result while cutting the reads and writes between HBM and on-chip SRAM. The memory did not “perform the computation”; the algorithm reduced the data movement needed to reach the same result. Optimization work in AI increasingly asks where data can be reused, not just how many arithmetic operations it takes.
11With multiple GPUs, the interconnect is the performance
Splitting a model or a batch across accelerators means moving gradients, activations, parameters, and token state between them. Fast chips sit and wait when the link between them is slow. Topology, collective communication, packaging, boards, racks, storage, power, and cooling are all part of the AI computer.
MLPerf compares complete systems against defined models, quality targets, scenarios, divisions, and versions. That is more reproducible than lining up TOPS figures reported under different precision and software assumptions, though no benchmark covers every workload.
12Quantization, offload, and distribution are not free speed-ups
Quantization cuts the bits per parameter and can save capacity and bandwidth, at the cost of accuracy, calibration work, kernel support, and metadata. Offloading to CPU memory or an SSD can make a larger model fit, at the cost of latency across PCIe or storage.
Sparsity, mixture-of-experts routing, distillation, and simply smaller models can all reduce the compute required, each with its own trade-offs in quality, routing, implementation, and task fit. The best system is not automatically “the largest GPU”; it is the one that meets the latency, throughput, quality, cost, energy, and data-governance requirements at hand.
13Editorial perspective: the future is heterogeneous and data-movement-aware
This section is our interpretation. We expect the next era of computing to lean further into heterogeneous systems rather than crown a single processor as the winner: CPUs for flexible control and latency-sensitive work, GPUs, NPUs, and ASICs for the parallel kernels that suit them, and memory hierarchies and interconnects for delivering the data.
Chiplets and fast links let functions be composed; near-memory and processing-in-memory research aims to cut movement; and any useful QPU is more likely to arrive as a co-processor for narrow algorithms. Competition will shift away from a chip’s peak FLOPS or TOPS and toward how efficiently a complete hardware and software system turns useful data into results with less movement and less energy. This is a scenario inferred from Roofline behavior, IO-aware AI work, chiplet standards, and the current state of quantum error correction, not a forecast of products or dates.
Primary sources
- Williams et al. — Roofline: An Insightful Visual Performance Model
- NVIDIA — CUDA C++ Programming Guide
- Jouppi et al. — In-Datacenter Performance Analysis of a Tensor Processing Unit
- Vaswani et al. — Attention Is All You Need
- Micikevicius et al. — Mixed Precision Training
- Dao et al. — FlashAttention
- Kwon et al. — Efficient Memory Management for Large Language Model Serving with PagedAttention
- Rajbhandari et al. — ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- MLCommons — MLPerf Benchmarks
- UCIe Consortium — Specifications
- Nature — Quantum error correction below the surface-code threshold
Read next
Can Quantum Computers Break Bitcoin? Qubits, Error Correction, and Cryptographic Migration11 min readRelated topics
Go deeper
Citation
- Title
- Why AI Needs GPUs and Memory: The Roles of CPUs, GPUs, and HBM
- Source
- Bitcoin Library (bitcoin.ne.jp)
- Canonical URL
- https://bitcoin.ne.jp/en/learn/ai-computing
- Author
- KK siiiiiixth
- Topic
- ai-computing
- Published
- Updated
- Last verified
- Editorial policy
- https://bitcoin.ne.jp/en/editorial-policy
- About
- https://bitcoin.ne.jp/en/about
- License
- Content reuse terms
Operator-owned article text, original diagrams, and public data may be used for citation, summarization, indexing, search, RAG, machine analysis, and AI model training. When content is presented to readers, identify Bitcoin Library and the applicable canonical URL where technically practicable.