Skip to content

Shelf 1 · Computing Foundations · 1 / 45

Computing Hardware: From von Neumann to the CPU, GPU, FPGA, and ASIC

From EDVAC and the transistor through Cell/B.E., CPUs, GPUs, FPGAs, and ASICs: a primary-source history of the PS3 Folding@home client, BOINC, Bitcoin mining, the companies, and the supply chain.

Check this article’s sources (60)

Article brief

Follow one chip through time and three machines join a single story: home PCs searching the sky, game consoles folding proteins, and hardware built only to mine Bitcoin.

A useful mental model

Picture a CPU as a few versatile cooks, a GPU as a large kitchen repeating work in parallel, an FPGA as a kitchen whose layout can be rewired, and an ASIC as an appliance perfected for one dish.

Where the analogy stops

Real performance also depends on memory traffic, software, the manufacturing process, power draw, and the workload itself. No architecture is fastest at everything.

You will be able to explain why the same idea of “computing power” suited BOINC yet took a very different shape in Bitcoin mining.

Open the glossary
Article contents14 chaptersJump to a chapter

1Computing history is not one upgrade path

Figure 1 Computing history is not a straight replacement of CPUs by GPUs and ASICs. Stored-program design, transistors, integrated circuits, MOS, microprocessors, FPGAs, GPUs, distributed science, and Bitcoin ASICs overlap, while older device classes persist in control, general computation, and verification roles.

The CPU, GPU, FPGA, and ASIC are not four generations in which each new device replaced the one before it. A CPU handles complex control and a wide range of software; a GPU pushes high throughput across many similar operations; an FPGA rewires its circuits in the field; an ASIC fixes its circuitry around a narrow task. Modern systems combine all of them with controllers, networks, and storage.

At least six axes matter: the switching medium (relay, vacuum tube, transistor), integration onto a chip, the instruction set architecture (ISA) and software compatibility, whether the circuit can be reprogrammed, how fine-grained the parallelism is, and the institutions that design, manufacture, and operate the machine. Transistor count, clock speed, core count, FLOPS, and hashes per second do not add up to one universal ranking.

ASIC is used here as a working term for an integrated circuit that is optimized for one domain and whose logic is not reconfigured in the field. In a broad sense, CPUs and GPUs are application-oriented chips too; the narrower usage is what lets us tell general processors, reconfigurable logic, and fixed-function accelerators apart.

2EDVAC and von Neumann: authorship versus a shared invention history

The First Draft of a Report on the EDVAC, issued in 1945 under John von Neumann’s name, described functional organs for arithmetic, central control, memory, input, and output. It put into wide circulation a logical design in which instructions and data alike could be held as numbers in electronic memory, and it became a central document for what was later called “von Neumann architecture.”

Von Neumann did not invent the stored-program computer on his own, however. The First Draft grew out of the Moore School collaboration around ENIAC and EDVAC, which included J. Presper Eckert, John Mauchly, Herman Goldstine, and others. Whose name appears on a report and who should be credited with the underlying invention are separate historical questions.

“Stored program” is itself a category historians applied later; the report of the day does not use it. Cutting down on rewiring, running code out of writable memory, holding instructions and data in the same memory, and the modern model of software are related claims, but they are not the same claim. Any “world first” therefore has to say what it is counting when it compares the Manchester Baby, the converted ENIAC, EDSAC, and other machines.

3From vacuum tubes to transistors, integrated circuits, and MOS

Vacuum-tube machines on the scale of ENIAC worked electronically, but they needed large numbers of tubes, a great deal of power and cooling, room for the wiring, and constant maintenance. In December 1947, John Bardeen and Walter Brattain demonstrated amplification in a germanium point-contact transistor at Bell Labs. William Shockley led the research group and later conceived the junction transistor. Their contributions were distinct and should not be collapsed into a single inventor.

Jack Kilby at Texas Instruments in 1958 and Robert Noyce at Fairchild Semiconductor in 1959 independently opened routes to the integrated circuit. Kilby’s monolithic idea and Noyce’s pairing of planar processing with metal interconnect were different contributions. In 1960, Mohamed Atalla and Dawon Kahng demonstrated the MOS transistor at Bell Labs, and dense digital circuits were built on it afterward.

Integration was not only miniaturization. Shorter wires, fewer connection points, repeatable photolithography, many dies per wafer, packaging, testing, and manufacturing yield together made complex processors reproducible at scale. The process-node labels quoted in nanometers today do not literally measure one universal physical feature.

4The microprocessor: putting a CPU on a chip, and carrying software across generations

The Intel 4004 grew out of Busicom’s calculator project and was commercialized in 1971 as a 4-bit CPU by a team that included Federico Faggin, Ted Hoff, Stanley Mazor, and Masatoshi Shima. It sits at the center of the history of early commercial single-chip microprocessors, but a definition that also counts the AL1, the MP944, or multi-chip designs gives a different answer to “first.” Intel did not invent the CPU as a concept, and no one person built the 4004 alone.

Processor history is also the history of instruction sets. IBM System/360 showed in 1964 what a compatible family of computers could be. The x86 line that began with Intel’s 8086 carried a large inheritance of software with it, and AMD64 extended x86 to 64 bits while leaving a path open for 32-bit software. RISC research and Arm opened up other implementation choices. Arm built its business for years on licensing architectures, processor IP, and compute subsystems, then announced its first Arm-designed production silicon, the AGI CPU, in March 2026, naming TSMC as its manufacturer. That does not mean Arm fabricates every Arm-based chip.

Moore’s 1965 article observed, and then projected, a rapid rise in the economically optimal density of components on a chip. It was not a law of nature, and it was not a promise that performance would double on a fixed schedule. More transistors still run into limits of memory latency, power density, and how much parallelism software can express, which pushed designs toward caches, SIMD, out-of-order execution, multicore processors, and accelerators.

5Compute units are not enough: memory, data movement, and parallelism

Figure 2 Attainable throughput is not determined by peak compute alone. At low arithmetic intensity, memory bandwidth and data movement impose the ceiling; only with enough data reuse can a workload approach the compute limit. This is a conceptual roofline relation, not a benchmark of any named CPU or GPU.

A CPU interprets instructions and works through complicated branches with an emphasis on latency. It contains several kinds of parallelism of its own: pipelines, superscalar execution, branch prediction, SIMD, and multiple cores. A GPU groups many threads together and hides waiting time with other work, which raises throughput. A CPU core and a GPU execution unit are not comparable things to count as “cores.”

Performance also depends on where the data sits (register, cache, main memory, VRAM, storage, or network) and on how many bytes have to move. Work with low arithmetic intensity becomes bandwidth-bound, and the cost of moving data to a GPU can exceed the time the GPU saves. Bitcoin’s SHA-256 repeats regular integer operations over a small state, which places it in a different region of the design space.

A benchmark is a comparison only when the application, input, precision, compiler, driver, memory capacity, wall-power boundary, and cooling conditions are the same. Peak FLOPS is not application speed, and hashes per second is not floating-point performance. A hardware record kept by an institution says which bottleneck a given number actually measured.

6The GPU: from graphics pipelines to general-purpose parallel computing

Graphics accelerators and dedicated graphics processors existed before 1999. That year NVIDIA marketed the GeForce 256 as the “world’s first GPU” under its own definition of a single-chip transform, lighting, triangle-setup, and rendering processor. Popularizing the GPU as a product category mattered, but it does not mean NVIDIA invented graphics acceleration.

Programmable shaders let general-purpose computing on GPUs move out from behind the graphics APIs. NVIDIA introduced CUDA in 2006, and Khronos released OpenCL 1.0 in 2008. In the usual heterogeneous model, the CPU still controls program flow, memory allocation, and kernel dispatch, while the GPU runs the highly parallel regions. The GPU does not replace an entire application on its own.

AMD’s graphics lineage runs through ATI, which it acquired in 2006; today AMD designs CPUs, GPUs, and adaptive computing products. Intel supplies integrated and discrete GPUs as well. Even when BOINC detects an NVIDIA, AMD, or Intel GPU, that device does science only if the project ships a compatible application and the driver environment it needs.

7FPGA and ASIC: reconfiguring a circuit or fixing it in place

Figure 3 CPU, GPU, FPGA, and ASIC differ in execution model and unit of change, not in one speed ranking. CPUs and GPUs change through software, FPGAs load circuit configurations, and ASICs require silicon redesign, producing different tradeoffs in generality, purpose-specific implementation NRE and redesign burden, production efficiency, and resilience to algorithm changes. Costs depend on the design and application rather than forming one price ranking.

An FPGA changes how its logic blocks and interconnect are configured after it has been deployed. Instead of stepping through software instructions, it can hold a purpose-built data path as an actual circuit. The Xilinx XC2064 of 1985 is an early commercial example. Design, synthesis, place-and-route, and timing closure all take work, but the circuit can still be updated after manufacture, which suits prototypes, smaller volumes, and low-latency pipelines.

An ASIC gives up that flexibility so it can drop the instruction decoding and general-purpose resources it does not need, concentrating area, throughput, and energy on one narrow job. The price is high design, verification, and mask costs, long lead times, and obsolescence as soon as the algorithm changes. An ASIC is not just a fast GPU, and an FPGA is not just a slow ASIC. What separates them is when each one locks its decisions in.

Crediting a company also requires dates. When Bitcoin FPGA miners appeared in 2011, Xilinx and Altera were independent companies. AMD acquired Xilinx in 2022. Intel acquired Altera in 2015, then sold a 51% interest to Silver Lake in September 2025 while keeping 49%. Buying a company later does not make its earlier inventions yours retroactively.

8Cell Broadband Engine: the heterogeneous multicore that turned PS3s into science accelerators

Figure 4 The shared word “Cell” hides two unrelated lineages. Cell Broadband Engine was a processor co-designed by Sony, Toshiba, and IBM; NTT DATA cell computing was a distributed-computing product and service brand. PS3 joined Folding@home’s own infrastructure, while Gene embedded BOINC to dispatch jobs across managed LAN PCs. Distinguish them by technical layer, operator, device ownership, and participation model.

Sony, Toshiba, and IBM began designing the Cell Broadband Engine (Cell/B.E.) together at the STI Design Center in Austin in 2001. Its Power Architecture-based PPE handles the operating system, branching, and work assignment, while several SIMD-oriented SPEs use local stores and DMA for data-parallel kernels. The general Cell design has eight SPEs, but scientific applications on the PlayStation 3 could use six. What a processor specification contains and what a product exposes are two different boundaries.

Stanford University’s Folding@home and Sony Computer Entertainment released a PlayStation 3 client in March 2007. A PS3 connected to Folding@home’s own client-server infrastructure, received work that fed molecular-dynamics trajectories for protein folding, misfolding, and related phenomena, computed it on Cell/B.E., and sent the output back. Folding@home was not a BOINC project, and the PS3’s RSX graphics processor was not the part doing the folding work. The compute engine was Cell/B.E.

A 2009 IPDPS paper reported roughly 50,000 active PS3s contributing about 1,400 TFLOPS at the time, within a conservative estimate of roughly 4,800 TFLOPS for the whole Folding@home system measured with real application code. Sony’s 2008 CSR report separately recorded more than 1.4 million cumulative PS3 participants and the petaflop milestone. Cumulative or registered participants, hosts active at once, theoretical peak, and application throughput are different measurements.

The peer-reviewed Cell molecular-dynamics work showed how SIMD execution, local stores, and DMA could speed up the simulation kernels that suited them, and it documented the limits as well: memory, data movement, and the narrow set of algorithms that mapped cleanly onto the hardware. A PS3 was not a general-purpose supercomputer. It substantially widened the research platform for the simulations it fitted, but no treatment or discovery should be credited to the PS3 alone. Trajectory analysis, hypothesis testing, replication, experiments, and review all come later.

The PS3 client ran until 6 November 2012. For that period it was a scientific instrument assembled out of uniform consumer hardware, worldwide deployment, Sony’s distribution channel, and Stanford’s research software. IBM and Toshiba co-designed Cell; the direct collaboration on the Folding@home client was between the Stanford team and Sony Computer Entertainment. Designing the chip, running the console platform, and stewarding a research project are separate contributions.

9BOINC: coordinating mixed hardware without flattening the differences

BOINC is middleware, not the science itself. It distributes project-specific applications to a wide range of hosts. A platform here is roughly a processor architecture paired with an operating system; a single research application may exist in CPU, GPU, OS-specific, and driver-specific versions. The scheduler uses plan classes and data about the host to judge compatibility, expected speed, how much CPU help a job needs, and which accelerator it wants.

GPUGrid used CUDA in 2008, BOINC announced NVIDIA GPU support, and SETI@home released a CUDA application that December. None of that made every science workload faster on a GPU. Regular kernels with high arithmetic intensity speed up well; heavy branching, large memory footprints, network I/O, porting cost, double precision, and short-lived research code can each stop a port. Keeping the CPU version is a decision about the workload and about who can take part, not a sign of falling behind.

Different hardware, compilers, and math libraries can change floating-point rounding and the order of operations. Two correct results, one from a CPU and one from a GPU, may not be bit-identical, so projects use tolerances, application-specific validators, homogeneous redundancy, or comparison within a single application version. BOINC handles that variety with metadata and validation rather than pretending it is not there.

NTT DATA’s cell computing Gene, released in 2005, has “Cell” in its name but has nothing to do with the Cell/B.E. processor. Gene combined BOINC, other open-source components, and NTT DATA software to distribute jobs across LAN PCs owned or managed by a single organization. A shared word is not evidence that Gene and the PS3 Folding@home client share a hardware lineage.

10Bitcoin mining: from CPUs to dedicated SHA-256 pipelines

Early Bitcoin clients included a CPU miner, and an SSE2 SHA-256 optimization was shared in 2010. A public OpenCL GPU miner followed the same year, open-source FPGA miners targeted Xilinx and Altera devices in 2011, and Canaan’s SEC filing records that its predecessor team shipped Avalon ASIC machines in January 2013. These transitions overlapped; they are not dates on which every participant switched at once.

SHA-256d has a fixed input structure and a fixed sequence of operations; mining varies the candidate block-header data over and over to find a hash below the target. GPUs ran many trials in parallel, FPGAs expressed the hash pipeline in configurable logic, and ASICs fixed the control and data paths around SHA-256. Because revenue tracks hash rate directly, set against electricity and capital costs, mining created a strong economic incentive to specialize.

Intel announced its own Blockscale SHA-256 ASIC in April 2022, quoting up to 580 GH/s and up to 26 J/TH at the chip boundary, not for a complete miner or a facility. In April 2023 Intel issued a product-discontinuance notice with a last-order date of October 20, 2023 and a forecast last-shipment date of April 20, 2024. How large an entrant is and how long it stays in the business are separate facts.

The whitepaper’s “one-CPU-one-vote” is a period expression for computing power, not a standing rule that gives every x86 CPU a vote. Chain selection compares cumulative proof of work. Neither the protocol nor a solo miner needs a central job server; a pool hands out jobs and measures lower-difficulty shares only when miners choose to work that way. Full nodes trust no vendor: they validate the block rules and the proof of work themselves.

11Shared silicon, different work and different validation

Comparison table for Shared silicon, different work and different validation
DimensionBOINC / scientific computingBitcoin mining
ProblemProject-specific models, searches, and analysisSearching for a hash below the target with SHA-256d
HardwareCPUs, GPUs, Arm devices, and more, chosen per applicationCompetitive mining now centers on SHA-256 ASICs
OutputResearch data, sometimes followed by analysis or experimentsA valid block candidate with its proof of work
ValidationReplication, tolerances, and project-specific validatorsEvery full node checks the hash and every block rule deterministically
MetricsApplication runtime, FLOPS, result instances, creditHash/s, J/TH, blocks, pool shares
IncentiveTaking part in research, credit, communitySubsidy, transaction fees, pool payouts

BOINC credit, WCG points, runtime years, FLOPS, and Bitcoin hashes per second do not convert into one another. Scientific FLOPS cannot predict SHA-256 hash rate, and a miner’s TH/s says nothing about scientific usefulness. What the two share is an institutional outline: hardware, electricity, cooling, networking, and software owned by many participants can be coordinated into computation on a planetary scale.

12The companies and the supply chain behind one chip

Figure 5 No computer or miner is the achievement of one company alone. Research and specifications, ISA and IP, chip design, EDA, lithography equipment, wafer fabrication, memory, packaging and test, boards, power, cooling, software, operation, and verification form a chain. Company contributions belong to specific layers and ownership periods rather than one undifferentiated “semiconductor industry.”

Intel designs x86 CPUs and other products and remains an integrated device manufacturer: its own fabs make most of what it sells. AMD designs x86 CPUs, GPUs, and the FPGA and adaptive products it inherited from Xilinx; NVIDIA designs GPUs and the CUDA software stack. AMD and NVIDIA depend mainly on outside foundries. TSMC is a pure-play foundry that manufactures customer designs; it does not design miners or run pools.

A working system can require instruction sets and IP, architecture, RTL, EDA tools, wafer fabrication, memory, packaging, testing, boards, power supplies, firmware, cooling, networks, and facility operations. Arm centers on architecture and IP while moving into Arm-designed silicon in 2026; ASML supplies lithography systems; Synopsys and Cadence, among others, supply EDA; TSMC, Samsung, and Intel fabricate wafers; OSAT firms such as ASE assemble and test. “Semiconductor company” therefore names no single role, and designing a chip does not mean owning the fab that makes it.

In Bitcoin ASICs, firms such as Canaan, Bitmain, and MicroBT design the chips and build the miner systems while contracting out parts of fabrication, packaging, and testing. The protocol, the firmware, the mining pool, the ASIC vendor, the foundry, and the operating facility are distinct actors. The same holds for BOINC: the Berkeley middleware, each research project, IBM’s or Krembil’s stewardship of WCG, the volunteer hosts, and the chip vendors should not be merged into one company’s achievement.

13Efficiency, electricity, and environmental claims need boundaries

A J/TH figure for an ASIC can refer to at least three boundaries: the chip, the miner measured at the wall, or the whole facility. A chip figure leaves out the power supply, fans, and controller; a miner figure leaves out transformation, pumps, and building cooling. Temperature, voltage, frequency, firmware, overclocking, and utilization all move the measurement. A product name on its own is not an efficiency record.

In BOINC, “using idle hardware” does not make the additional electricity zero. The variables you need are marginal wall power above baseline, runtime, successful completions, retries, cooling, and electricity specific to a place and a time. Credit, FLOPS, or total runtime on their own cannot give you kWh or carbon emissions.

Electricity used in operation also has to be kept separate from the lifecycle effects of wafer fabrication, packaging, transport, facility construction, and disposal. Better performance per watt does not guarantee lower total electricity use when network hash rate or participation grows. Efficiency, total consumption, carbon, and electronic waste are four different questions.

14How a computer history museum should read the evidence

Every “first” needs three qualifiers: first at what, under whose definition, and on whose authority. Vendor timelines, product documents, and SEC filings are primary evidence of what a company did, but they are also how a company describes itself. Read them alongside museum collections, original papers, and source code, so that a marketing first stays distinct from a historical attribution.

Do not credit the later parent with achievements that predate an acquisition. ATI’s graphics history joins AMD’s corporate history in 2006, and Xilinx joins AMD in 2022. Supported operating systems, drivers, project status, and corporate ownership are living evidence that carries a lastVerified date; a 1945 report or a 2013 paper is a fixed historical record.

Preserving a machine means more than displaying its case. It means holding together the circuits, the source, compilers, drivers, protocols, benchmark conditions, electricity boundaries, organizations, failed designs, and what participants actually did. Only then can we check how desktop CPUs and GPUs became scientific instruments through SETI@home and BOINC, and how a different set of conditions pushed Bitcoin toward a specialized ASIC industry.

Primary sources

Read next

Supercomputing: Its History, Role, and the Next Five and Ten Years12 min read
Share

Citation

Title
Computing Hardware: From von Neumann to the CPU, GPU, FPGA, and ASIC
Source
Bitcoin Library (bitcoin.ne.jp)
Canonical URL
https://bitcoin.ne.jp/en/learn/computing-hardware
Author
KK siiiiiixth
Topic
computing-hardware
Published
Updated
Last verified
Editorial policy
https://bitcoin.ne.jp/en/editorial-policy
About
https://bitcoin.ne.jp/en/about
License
Content reuse terms

Operator-owned article text, original diagrams, and public data may be used for citation, summarization, indexing, search, RAG, machine analysis, and AI model training. When content is presented to readers, identify Bitcoin Library and the applicable canonical URL where technically practicable.