## Why AI Needs GPUs and Memory: The Roles of CPUs, GPUs, and HBM

URL: https://bitcoin.ne.jp/en/learn/ai-computing
URL role: Self-canonical for this English representation
Japanese counterpart URL: https://bitcoin.ne.jp/learn/ai-computing
Slug: ai-computing
Language: en
First published: 2026-08-23
Last updated: 2026-09-03
Last verified: 2026-08-23
Description: How CPU latency, GPU throughput, RAM, VRAM, HBM, training state, inference, and the KV cache shape AI computers, and where computing goes next.

Ranking GPUs above CPUs misses how an AI computer actually works. The CPU handles control and low-latency response, GPUs and tensor accelerators supply regular parallel throughput, and memory holds the model and its intermediate state and feeds them to the arithmetic units. Performance depends on more than the arithmetic units: where the data sits, how many bytes move, which precision is used, and how the devices are connected all matter.

Reuse terms: https://bitcoin.ne.jp/licenses#content-reuse
Third-party sources retain their own rights. Cite the canonical HTML URL or its section fragment.
Sources below are article-level references; they do not establish support for every sentence. Check the original source and its date before quoting time-sensitive claims.
This is educational material, not investment or legal advice. Dates describe the editorial text; this representation excludes browser tools and live market/network values.

### A 30-second entry

Behind a single short reply to a prompt, CPUs, GPUs, memory, and networks each do a different job while moving enormous amounts of data.

Mental model: Picture the CPU as a stage manager making control decisions, the GPU as a large ensemble repeating structured work in parallel, memory as the worktables holding scores and intermediate results, and interconnects as the transport routes.
Where the analogy stops: This analogy explains roles, not a universal ranking. Real performance changes with the model, numeric precision, batch size, software, memory capacity, bandwidth and latency, communication, and power. Adding GPUs or memory does not always make a workload faster.
Reading journey: [Roles, not a ranking](https://bitcoin.ne.jp/en/learn/ai-computing#roles-not-ranking) → [Data movement before arithmetic](https://bitcoin.ne.jp/en/learn/ai-computing#data-movement) → [A heterogeneous future](https://bitcoin.ne.jp/en/learn/ai-computing#editorial-view)
What you will understand: You will be able to say what CPUs, GPUs, and memory are each good at without ranking them on one speed table, and locate an AI bottleneck across the whole system.

### The 30-second answer: CPUs, GPUs, and memory play different roles

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#key-facts

| Element | Strength | Typical AI-system role |

| --- | --- | --- |

| CPU | Low latency on a few threads, complex branching, broad compatibility | OS, I/O, preprocessing, scheduling, accelerator control |

| GPU / AI accelerator | High throughput across many parallel operations | Matrix and tensor work, training, batched inference |

| Memory | A hierarchy of capacity, bandwidth, and latency | Weights, activations, gradients, optimizer state, and KV cache |

| Interconnect | Moves data between devices | Synchronizing and distributing work across accelerators |

A GPU is not always faster than a CPU, and more memory does not make a model smarter. Any comparison has to weigh the workload, the software, the numerical precision, and the data movement together.

### The shape of the work matters more than a performance ranking

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#roles-not-ranking

Latency is how long one job takes to finish; throughput is how much work completes per unit of time. A CPU is built for a fast response along one or a few complex paths. A GPU is built for the combined rate of many similar ones.

Small inputs, heavy branching, sequential dependencies, or transfers between host and device can leave GPU lanes idle. When the same large matrices are processed over and over, many GPU lanes together can outrun a handful of powerful CPU cores. The first question is not which processor is “better,” but how regularly the work divides.

### CPU strengths: latency, branching, and control

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#cpu-strengths

A modern high-performance CPU core combines branch prediction, out-of-order execution, large caches, and a broad instruction set to handle shifting control flow with low latency. That suits responsive work: the operating system, networking, storage, databases, compression, tokenization, data loaders, and error handling.

CPUs also have SIMD units, matrix extensions, and multiple cores, so they can run AI inference themselves. That is a reasonable choice when the model is small, the batch is tiny, the latency target is strict, or moving data to an accelerator costs more than it saves. Even in accelerator-heavy systems, the CPU still prepares jobs, launches kernels, manages I/O, and handles failures.

### GPU strengths: parallelism and throughput

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#gpu-strengths

A GPU spreads one kernel across many lightweight threads to sustain high throughput on regular arithmetic. The model grew out of pixel and vertex work, and it maps well onto matrix and tensor operations over wide data. Dedicated matrix units can drop to lower precision to run the multiply-accumulate patterns that dominate AI.

Branch divergence, tiny jobs, frequent round trips to the CPU, or memory that cannot keep up will leave those lanes waiting. Hardware parallelism is usable only when compilers, libraries, and toolchains such as CUDA support it. Peak FLOPS or TOPS alone cannot tell you how an application will perform.

### Why AI uses so much matrix math

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#matrix-ai

Many neural-network layers multiply an input vector by a weight matrix and then apply a nonlinear transformation. In Transformers, both attention and the feed-forward networks contain large matrix multiplications. Because the same multiply-accumulate structure repeats, the work fits the parallel units in GPUs and TPUs.

AI is not just matrix multiplication. Normalization, activation, sampling, routing, preprocessing, and communication all matter. Training behaves differently from inference, and prompt prefill behaves differently from token-by-token decode. A benchmark is meaningful only when the model, the precision, the batch size, the quality target, and the scenario line up.

### Memory: keep capacity, bandwidth, and latency separate

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#memory-three

In conventional architectures, memory does not do the arithmetic. It stores data and delivers it to the units that do. Capacity decides what fits at once, bandwidth decides how many bytes move per second, and latency decides how long a single request waits. The three are independent: a large pool can be slow, and wide bandwidth does not remove random-access latency.

| Layer | Location and role |

| --- | --- |

| Registers / on-chip SRAM / cache | Small, close to arithmetic, and fast |

| VRAM / HBM | Large accelerator-local memory; HBM uses stacking and a wide interface for bandwidth |

| System DRAM | CPU-side main memory with capacity and generality |

| SSD / storage | Larger and slower than execution memory; used for loading and offload |

### Training must store more than weights

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#training-memory

Training holds gradients, optimizer state, and forward-pass activations on top of the model parameters. How much that adds depends on the precision and the optimizer, so multiplying the parameter count by a byte size understates it. ZeRO partitions optimizer state, gradients, and parameters across devices for exactly this reason: the memory bill constrains the design of the whole system.

Mixed precision computes and stores some values at lower precision, which improves throughput and effective capacity but calls for techniques that keep the numerics stable. Activation checkpointing recomputes selected activations instead of keeping them, trading compute for memory. Model parallelism raises the total capacity available but adds communication and synchronization.

### Inference holds weights and a growing KV cache

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#inference-memory

LLM inference reads the model weights again and again as it generates tokens. Autoregressive decode may do little new arithmetic per step, so it runs up against bandwidth rather than compute. Prompt prefill and larger batches can have higher arithmetic intensity, so “all AI inference is memory-bound” is too broad.

The KV cache keeps attention keys and values rather than recomputing them, and it grows with the number of layers, tokens, and concurrent requests. PagedAttention proposed managing that cache the way an operating system manages virtual memory, cutting the waste in contiguous allocation. Long context is therefore a question of memory capacity and bandwidth, not only of algorithmic capability.

### How many gigabytes does it take just to hold 7B or 70B weights?

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#capacity-arithmetic

Counting weights alone at two bytes per FP16 or BF16 parameter gives about 14 GB for 7 billion parameters and 140 GB for 70 billion. A theoretical packed 4-bit weight body would come to roughly 3.5 GB and 35 GB. These are decimal gigabytes.

Real systems then add quantization metadata, values kept at higher precision, runtime and allocator headroom, temporary buffers, and the KV cache. Training adds gradients, optimizer state, and activations on top. The arithmetic does not tell you how many GPUs to buy; it is a floor that shows why capacity decides where a model can run.

### Fast arithmetic still waits when the data cannot arrive

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#data-movement

The Roofline model bounds performance by the smaller of two quantities: peak compute, and memory bandwidth multiplied by operational intensity, the operations performed per byte moved. A kernel that moves a lot of data without much arithmetic hits the bandwidth roof, and adding arithmetic units does not help.

FlashAttention returns the same attention result while cutting the reads and writes between HBM and on-chip SRAM. The memory did not “perform the computation”; the algorithm reduced the data movement needed to reach the same result. Optimization work in AI increasingly asks where data can be reused, not just how many arithmetic operations it takes.

[Figure: ai-computing-system — An AI computer is not a GPU alone. CPUs handle control and I/O, GPUs, NPUs, and ASICs run suitable parallel kernels, and caches, SRAM, HBM, and DRAM hold and feed data. Interconnects, storage, software, power, and cooling complete the system, whose useful performance can be constrained by a slow path or insufficient capacity. This is a conceptual role stack, not a product speed ranking.]

### With multiple GPUs, the interconnect is the performance

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#multi-accelerator

Splitting a model or a batch across accelerators means moving gradients, activations, parameters, and token state between them. Fast chips sit and wait when the link between them is slow. Topology, collective communication, packaging, boards, racks, storage, power, and cooling are all part of the AI computer.

MLPerf compares complete systems against defined models, quality targets, scenarios, divisions, and versions. That is more reproducible than lining up TOPS figures reported under different precision and software assumptions, though no benchmark covers every workload.

### Quantization, offload, and distribution are not free speed-ups

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#tradeoffs

Quantization cuts the bits per parameter and can save capacity and bandwidth, at the cost of accuracy, calibration work, kernel support, and metadata. Offloading to CPU memory or an SSD can make a larger model fit, at the cost of latency across PCIe or storage.

Sparsity, mixture-of-experts routing, distillation, and simply smaller models can all reduce the compute required, each with its own trade-offs in quality, routing, implementation, and task fit. The best system is not automatically “the largest GPU”; it is the one that meets the latency, throughput, quality, cost, energy, and data-governance requirements at hand.

### Editorial perspective: the future is heterogeneous and data-movement-aware

Section URL: https://bitcoin.ne.jp/en/learn/ai-computing#editorial-view

This section is our interpretation. We expect the next era of computing to lean further into heterogeneous systems rather than crown a single processor as the winner: CPUs for flexible control and latency-sensitive work, GPUs, NPUs, and ASICs for the parallel kernels that suit them, and memory hierarchies and interconnects for delivering the data.

Chiplets and fast links let functions be composed; near-memory and processing-in-memory research aims to cut movement; and any useful QPU is more likely to arrive as a co-processor for narrow algorithms. Competition will shift away from a chip’s peak FLOPS or TOPS and toward how efficiently a complete hardware and software system turns useful data into results with less movement and less energy. This is a scenario inferred from Roofline behavior, IO-aware AI work, chiplet standards, and the current state of quantum error correction, not a forecast of products or dates.

### Primary Sources

- Williams et al. — Roofline: An Insightful Visual Performance Model: https://doi.org/10.1145/1498765.1498785
- NVIDIA — CUDA C++ Programming Guide: https://docs.nvidia.com/cuda/cuda-programming-guide/index.html
- Jouppi et al. — In-Datacenter Performance Analysis of a Tensor Processing Unit: https://arxiv.org/abs/1704.04760
- Vaswani et al. — Attention Is All You Need: https://proceedings.neurips.cc/paper/7181-attention-is-all-you-need.pdf
- Micikevicius et al. — Mixed Precision Training: https://openreview.net/forum?id=r1gs9JgRZ
- Dao et al. — FlashAttention: https://proceedings.neurips.cc/paper/2022/hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html
- Kwon et al. — Efficient Memory Management for Large Language Model Serving with PagedAttention: https://doi.org/10.1145/3600006.3613165
- Rajbhandari et al. — ZeRO: Memory Optimizations Toward Training Trillion Parameter Models: https://doi.org/10.1109/SC41405.2020.00024
- MLCommons — MLPerf Benchmarks: https://mlcommons.org/benchmarks/
- UCIe Consortium — Specifications: https://www.uciexpress.org/specifications
- Nature — Quantum error correction below the surface-code threshold: https://www.nature.com/articles/s41586-024-08449-y
