AI Inference Systems: Chip to Cluster · The machine
The machine you are learning to build
One request, one machine. Orbit it, open any part, and keep going until you reach the blocks on a single die — every part tells you what it does while a token is served, which numbers it has to meet, and which chapter trains you to build it.
Now inside Data hall row, 5 kinds of part.
Inside Data hall row
Prefill: reading your prompt. Your whole prompt goes through the machine once, in parallel. It is the arithmetic-heavy half of serving, and it ends with the first token you see. Press Step or Play to follow it.
Six levels, and what you build at each
Cluster
Data hall row
Read the workload first: two request traces become bytes per token, a first-token target and the tokens per second a row of racks has to hold.
Rack
Accelerator rack
Then the scale-up domain: how seventy-two accelerators are made to behave as one machine, and what its fabric, power shelves and coolant loop have to carry.
Tray
Compute tray
Then the sled you can pull out: the processors that schedule the work, the network cards that bring requests in and carry tokens out, and the cold plates over all of it.
Package
Accelerator package
Then the accelerator itself: a compute die, the stacks of high-bandwidth memory beside it, and the substrate that gets their signals across without losing them.
Die
Compute die
Then the floorplan: matrix engines, attention engines, scratchpad memory and the network on chip that joins them, placed where the data has to flow.
Block
Matrix engine tile
And then one block, all the way down: arithmetic, synthesizable design, verification, timing and a signed-off layout with the evidence that proves it.
