// -->

Thread Rating:
  • 0 Vote(s) - 0 Average
  • 1
  • 2
  • 3
  • 4
  • 5
The Argument to Rethink CPU Design Completely
#1
The Argument to Rethink CPU Design Completely

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

We've been building processors backwards for sixty years. Not wrong in the sense that they don't work — they work. But wrong in the sense that a tool designed without its job in mind will always be outperformed by one designed for it. The CPU is a tool designed in a vacuum, handed to programmers, and accompanied by a single instruction: figure out how to express your work in terms of what this thing can do.

That's not engineering. That's a workaround masquerading as a paradigm.

─── ◆ ───

The GPU Proved the Case

The GPU didn't emerge because someone had a cool idea about parallel processing. It emerged because the CPU failed. Real-time 3D graphics demanded a type of computation — the same floating-point operation applied to millions of vertices and pixels every sixteen milliseconds — that the CPU was structurally incapable of delivering. Not because the transistors were too slow. Not because the process node was too large. Because the architecture was wrong.

The response wasn't to make the CPU better at graphics. The response was to build a different machine entirely. And critically, it wasn't built by CPU engineers. It was built by graphics people — Jim Clark, Jensen Huang, the SGI lineage — people who understood the work and designed hardware to match it. They didn't start from "how do we improve the pipeline." They started from "what does this computation actually look like?" and built silicon that mirrored it.

Same transistors. Same silicon. Same fabrication process. Completely different chip. Orders of magnitude better for the workload it was designed around.

That's not an incremental improvement. That's an indictment.

─── ◆ ───

The Crack That Everybody Ignored

The GPU was the first crack in the CPU's claim to universality. The CPU world looked at it and categorized it as a "special-purpose accelerator for games." A peripheral. A toy. They went back to tweaking branch predictors and adding pipeline stages.

Then it happened again.

Google built the TPU because neither CPUs nor GPUs were actually right for inference. Another topology. Another order-of-magnitude gain. Another admission that the general-purpose processor couldn't do the job.

Then the NPU. Apple, Qualcomm, and everyone else bolting neural engines onto their SoCs. Another concession.

Then the DSP, the video encoder, the cryptographic accelerator, the image signal processor. Each one a specialized unit designed from the workload backward, each one an implicit admission that the CPU core — the von Neumann heart of the chip — is wrong for yet another class of computation.

The pattern screams at us. GPU: the CPU can't do graphics. TPU: the CPU can't do inference. NPU: fine, we'll glue a neural engine onto the side. DSP: the CPU can't do signal processing fast enough. Every single one is the industry saying "this architecture is wrong for this work" and then, instead of rethinking the architecture, building a new thing next to it and leaving the CPU untouched.

At some point you have to ask: if the CPU needs a co-processor for graphics, a co-processor for AI, a co-processor for signal processing, a co-processor for video, and a co-processor for cryptography — what is the CPU actually still good at? What workload is it the right architecture for?

The honest answer: running legacy software and managing control flow for irregular serial tasks. That's what sixty years of optimization produced. A very expensive, very power-hungry traffic cop that excels at running an OS kernel and a Python interpreter and not much else. Everything computationally significant has migrated to specialized hardware designed workload-first.

─── ◆ ───

The Von Neumann Bottleneck Is the Symptom, Not the Disease

Every processor built today — CPU, GPU, all of them — is a von Neumann machine at its core. Compute happens here, data lives over there, and a bus moves data back and forth between them. The entire memory hierarchy exists because of this separation: registers, L1, L2, L3, DRAM, storage. Each level is a coping mechanism for the fundamental problem that compute and data are in different places.

The numbers tell the story. A 64-bit floating-point multiply in a modern process costs roughly one to two picojoules. Moving that same 64-bit value from DRAM to the compute unit costs ten to twenty nanojoules — roughly ten thousand times more energy. Even pulling it from L1 cache costs around fifty times the energy of the computation itself.

The machine we built burns 99% of its energy on logistics and 1% on work. That's not a computer. That's a trucking company that occasionally does arithmetic.

Cache is a bandage. Prefetchers are a bandage on the bandage. The entire memory hierarchy is an elaborate mitigation strategy for a decision made in 1945: put the arithmetic unit in one place and the memory in another. Every generation we make the cache bigger, add another level, make the prefetcher smarter — spending more transistors managing data movement instead of questioning why data moves at all.

But the von Neumann bottleneck is a symptom. The disease is deeper.

─── ◆ ───

Building Tools Backwards

Here's the core of the problem: we designed a processor and then asked programmers to fit their work to it. That's backwards.

A tool should be designed for the work it needs to do. You don't build a hammer and then go looking for things to hit. You analyze the job, understand the forces and materials involved, and design a tool that matches them. The CPU was designed around what was convenient to build in the 1960s — a sequential instruction stream, a centralized register file, a single program counter — and then the entire software world was told to express all computation in those terms.

It worked, for a while. Early computing really was sequential: solve a differential equation, sort a payroll file, evaluate a logical expression. Single-threaded, branchy, serial. The von Neumann architecture was a reasonable match for that workload profile.

But workloads evolved. Graphics, simulation, networking, databases, machine learning, data analytics, genomics — each one has a fundamentally different computational structure. None of them look like a sequential instruction stream with unpredictable branches. But the CPU didn't evolve to match. It kept the same basic topology and tried to make it faster: deeper pipelines, wider issue, out-of-order execution, speculative execution, bigger caches. Billions of transistors spent maintaining the illusion of a fast sequential machine while the actual work became increasingly parallel, regular, and data-dominated.

The GPU succeeded because it was built the right way. The work came first. The question "what does rendering actually look like?" produced an architecture — thousands of simple cores, massive memory bandwidth, deep thread parallelism — that was derived from the computation. The hardware mirrored the work.

And when AI arrived with workloads that looked structurally identical to rendering — enormous regular matrix operations on streaming data — the GPU was accidentally perfect for it. Not because NVIDIA anticipated AI, but because they'd built hardware shaped by computational patterns rather than by architectural tradition.

─── ◆ ───

The Comfort Trap

The CPU world watched the GPU revolution happen from the sidelines and did nothing. Not because they lacked talent — Intel had the best process engineers on Earth, the most advanced fabs, and virtually unlimited capital. They did nothing because they were comfortable. The existing paradigm worked. It made money. Each generation delivered a measurable improvement. Why rethink the foundation when the foundation still pays?

This is the deadliest trap in engineering: the incremental gain that confirms the paradigm. Every 10–15% IPC improvement reinforces the belief that the approach is sound, that optimization within the current framework is the right strategy. The gains are real. The products ship. The revenue comes in. And the structural problem that would require a fundamental rethink gets buried under evidence that the current path is "working."

It is working. In the same sense that putting a bigger engine in a horse-drawn carriage is working. You're going faster. You're also optimizing the wrong machine.

The GPU broke out of this trap because it had the advantage of having nothing to protect. There was no legacy graphics codebase demanding backward compatibility with a von Neumann model. NVIDIA could design the hardware to match the work, write new programming models from scratch (CUDA, shader languages), and build a coherent stack with no baggage. The CPU world can't do that. Every improvement to a CPU must be backward-compatible with x86 or ARM. The instruction set, the memory model, the sequential execution contract — all of it must be preserved. The result is a chip that spends billions of transistors pretending to be a fast PDP-11 while the world it's serving looks nothing like what a PDP-11 was designed for.

─── ◆ ───

What the Right Answer Actually Looks Like

If we take the principle seriously — design the hardware from the work, not the other way around — several things follow immediately.

Compute and memory should be the same chip. Not "add a bigger cache." Not "put HBM next to the GPU." Actually embed arithmetic capability inside the memory fabric so that data never moves. Apply voltages to rows, let currents through resistive elements perform multiplication, read results on columns. One operation. Zero data movement. The physics of this works today. Companies have built functional prototypes. The bottleneck is the software ecosystem that assumes von Neumann, not the silicon.

The instruction stream is the wrong abstraction. Fetching, decoding, and executing instructions one at a time — even out of order, even speculatively — is a bizarre way to compute when the work is "multiply these two enormous matrices" or "apply this filter to a billion records." The computation should be expressed as a dataflow graph and mapped spatially onto silicon. Data enters, flows through functional units, results emerge. No program counter. No branch predictor. No instruction cache. The chip is the computation.

The concept of a single general-purpose processor is itself the problem. Instead of one chip that does everything adequately, design purpose-matched silicon for the dominant computational patterns of the era. Not as bolted-on accelerators managed by a von Neumann traffic cop — as first-class compute substrates, each with its own memory, its own data model, and its own programming interface. The GPU already proved this works for one workload class. Extend the principle.

The software stack has to change. This is the part nobody wants to hear. The programming models, the compilers, the operating systems, the languages — all built around sequential execution on von Neumann hardware — must change. The GPU proved this is survivable. CUDA didn't exist before the hardware demanded it. Programmers learned it because the performance advantage was undeniable. The same will happen again when hardware built from the workload backward delivers not 15% but 100x gains for the work that matters.

─── ◆ ───

The Opportunity

The physics is ready. We've spent fifty years perfecting silicon — crystal growth, thermal oxidation, photolithography, etching, doping, metallization. The process chain from ingot to packaged chip is the most refined manufacturing discipline in human history. The transistors are extraordinary. We can build gate-all-around nanosheet structures at 2nm, stack billions of devices on a single die, and achieve switching energies approaching fundamental thermodynamic limits.

None of that needs to change. The materials are right. The fabrication is right. The transistors are right. What's wrong is how we organize them. The architecture, the topology, the fundamental conception of what a processor is and how it relates to the work it's meant to do — that's where the breakthrough lives.

The GPU proved that reorganizing the same transistors around the actual structure of a workload can produce order-of-magnitude gains. The TPU proved it again. The NPU proved it again. Every specialized accelerator ever designed proved it again. The evidence is overwhelming and the lesson is clear: match the silicon to the work, not the work to the silicon.

The company that fully internalizes this — that builds a processor from the workload backward with no loyalty to von Neumann, no backward compatibility constraints, and no institutional comfort with the status quo — will do to the CPU what the GPU did to fixed-function graphics pipelines. It won't be a generational improvement. It will be a category reset.

The trillion dollars invested in the current paradigm is not an argument for continuing it. It's the weight of the trap. Getting comfortable with what exists is the most expensive mistake in engineering — not because it doesn't work, but because it prevents you from seeing what would work better.

Fifty years of geometry changes on the same basic machine. The next fifty should be about building a fundamentally different machine from the same excellent geometry. The transistors are waiting. The architecture hasn't caught up.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
— Z E R O S  T O  H E A V E N ! —
Photonamus Industries • Founder
Reply


Forum Jump:


Users browsing this thread: 1 Guest(s)