AI Inference Systems: Chip to Cluster · The machine

The machine you are learning to build

One request, one machine. Orbit it, open any part, and keep going until you reach the blocks on a single die — every part tells you what it does while a token is served, which numbers it has to meet, and which chapter trains you to build it.

Your system asks for less motion, so the path is followed one step at a time with Step.Choose a part inside this one first, in the model or in the list.You are looking at the whole machine.The path is at its last step. Reset it to follow it again.The path has not started.

Now inside Data hall row, 5 kinds of part.

Drag or scroll to move down the page · Alt + drag to orbit · Alt + scroll to zoom · double-click a part to open itDrag sideways to turn · pinch to zoom · double-tap a part to open it. Arrow keys orbit, Control and an arrow move the model across the frame, plus and minus zoom, 0 resets the view, Space chooses a part and Enter opens it.

Inside Data hall row

Prefill: reading your prompt. Your whole prompt goes through the machine once, in parallel. It is the arithmetic-heavy half of serving, and it ends with the first token you see. Press Step or Play to follow it.

Six levels, and what you build at each

  1. Cluster

    Data hall row

    Read the workload first: two request traces become bytes per token, a first-token target and the tokens per second a row of racks has to hold.

  2. Rack

    Accelerator rack

    Then the scale-up domain: how seventy-two accelerators are made to behave as one machine, and what its fabric, power shelves and coolant loop have to carry.

  3. Tray

    Compute tray

    Then the sled you can pull out: the processors that schedule the work, the network cards that bring requests in and carry tokens out, and the cold plates over all of it.

  4. Package

    Accelerator package

    Then the accelerator itself: a compute die, the stacks of high-bandwidth memory beside it, and the substrate that gets their signals across without losing them.

  5. Die

    Compute die

    Then the floorplan: matrix engines, attention engines, scratchpad memory and the network on chip that joins them, placed where the data has to flow.

  6. Block

    Matrix engine tile

    And then one block, all the way down: arithmetic, synthesizable design, verification, timing and a signed-off layout with the evidence that proves it.

Dissect the machine — AI Inference Systems: Chip to Cluster — ActGen