Learn / AI Inference Systems: Chip to Cluster
AI Inference Systems: Chip to Cluster
Training discussion board — every chapter's threads and your study groups, in one feed.
Start with short interactive and long-context requests, then design the system that serves them: prefill, KV state and autoregressive decode. Build a compute, memory, interconnect or control component; verify its numerical and physical behavior, integrate analog interfaces and software, and carry its limits through chip, package, host, rack and cluster design. Compare quality, TTFT/TPOT and useful output tokens per facility energy and cost, then publish a reproducible portfolio.
Prerequisites: Basic digital logic and circuits: flip-flops, transistors and op-amps. No previous EDA-tool or tapeout experience is required. Reading and chapter checks are free. Workspace and AI work use credits; review the displayed price and runtime limit before starting each practical, and reuse the project you are developing.
This page is open to everyone. The chapters open once you sign in — free, with GitHub — because your progress, your runs, and your credential all belong to an account.
1 of 6
Seventy-two accelerators. One machine.
A row of liquid-cooled racks, a network spine and the power and cooling that let them answer one request together. Every token you are served leaves a room like this.
2 of 6
Eighteen trays, one fabric.
Inside a rack, compute trays and switch trays share a copper backplane, so seventy-two accelerators behave like one very large chip. Power shelves and a coolant loop keep it alive.
3 of 6
Four packages, two hosts, one sled.
A tray carries the accelerators, the processors that schedule their work and the network cards that bring requests in and carry tokens out, all under cold plates.
4 of 6
Eight stacks of memory around one die.
The weights and the KV cache live in high-bandwidth memory stacked beside the die, close enough that a decode step can stream them every few milliseconds.
5 of 6
Where every multiply happens.
Matrix engines, attention engines, scratchpad memory and a network on chip: the blocks a decode step runs through, laid out on one piece of silicon.
6 of 6
Now build one of these.
Every chapter of this training trains you to design, verify and sign off one of these parts. Pick a part, open its chapter, and start.
Chapter 1 of 36
What an Inference Part Has to Move
Read the lesson, apply it to your project, check your understanding, then continue to the next chapter.
Explore any chapter, even before finishing the previous one.
Technical interview: voice or text
About an hour with an AI interviewer, on the track you choose. Type your answers, or use spoken questions and dictation where your browser supports them. You can edit the answer before submitting it. A pass stands for a year. It is what opens paid work on the platform: the signing bench, contracts, and expert matching.
The credential is the training’s proof of the work; the interview is where you defend it. You can sit it now or after the credential. You can start or continue the course without it.
Take the technical interviewChoose the component that changes your serving system
Start with your short and long-context request traces. Choose one component to implement, then use its results to revise the chip, package, host, rack and cluster design in the same project repository.
Sign in to save this choice across your laptop and phone. You can explore the syllabus first.
Use the same serving scorecard
Compare designs on the same request distribution, concurrency, cache locality and output-quality target. Report TTFT and TPOT with tail latency, and count useful output tokens that meet those targets per facility joule and per unit cost. Count energy and cost for compute, memory, hosts, networking and cooling, including idle time and rejected work.
- 1. Turn request traces into a hardware mission
Two request traces, quality and TTFT/TPOT targets, a component choice, and a first silicon run that establishes your evidence workflow.
- 2. Build the compute and KV-state path
A checked numerical reference, compute/memory interfaces and verified state transitions for prefill and decode, with accepted-token and rollback tests for the optional draft-and-verify path.
- 3. Prove the component at a physical operating point
Timing, area, power, reset and signoff evidence that replaces an assumed component rate in your serving model and exposes the next bottleneck.
- 4. Supply the chip with data, power and cooling
Package and board interfaces, an analog sensor/amplifier and mixed-signal checks, plus delivery and thermal limits for bursty prefill and sustained decode.
- 5. Serve the workload across hosts, racks and clusters
A software path and placement/scheduling model that reconciles request concurrency, cache locality, network traffic and failures with useful tokens per facility energy and cost.
- 6. Release and defend the complete inference design
A reproducible component, a chip-to-cluster dossier and a release plan that let another engineer rerun the workload comparison and inspect every engineering claim.
Your first mission
Make the next token arrive on time
Design a serving system for a mix of short interactive requests and long-context requests. Follow each request through prefill, KV state and autoregressive decode; build one component and use its measured limits to size the chip, package, host, rack and cluster.
Create two illustrative request traces using the same model and tokenizer: a short conversation with 2,048 input and 256 output tokens, and a long-context conversation with 32,768 input and 128 output tokens. Compare cold and warm prefixes at 1, 8 and 32 concurrent requests. Record arrival times and set the same output-quality, time-to-first-token (TTFT) and time-per-output-token (TPOT) gates before choosing hardware.
1. Prefill the prompt
Trace matrix work and memory traffic until the first token. Separate reusable prefix state from work that must run again.
2. Keep the right KV state
Track which request owns each cache page, how context grows, and what moves when requests share, migrate or release state.
3. Decode and serve
Generate successive tokens under concurrent arrivals. As an extension, add draft-and-verify decoding: count accepted tokens, rejection work and KV rollback while preserving the quality target.
Compare request execution in the laboratory below, then test a first memory and power estimate. Save your assumptions in the project workbook, choose a component to build, and revise that estimate as your experiments produce evidence.
Watch the mechanism
From prompt to cluster
Follow prompt processing, cached state and decoding, then see what servers and a connected cluster must supply.
0:42 · Narrated video · Optional background music · Play, pause, or seek at your own pace.
Use fullscreen for a closer view, or read the transcript below.
Keyboard shortcuts
Tab to the video area, then use these keys.
- K / Space
- Play or pause
- J / L
- Back or forward 10 seconds
- ← / →
- Back or forward 5 seconds
- ↑ / ↓
- Raise or lower volume
- M / C
- Mute or captions
- F / I
- Fullscreen or miniplayer
- Home / End
- Jump to the start or end
- 0–9
- Jump to a tenth of the video
- Shift + , / .
- Slower or faster playback
- , / .
- Step back or forward while paused (about one frame)
Read the full transcript
- Every token takes a physical journey. Process the known prompt under a causal mask.
- Store keys and values. This cached state and model weights feed decoding.
- Repeat attention and feed-forward layers. After the final layer, select a token.
- Check optional drafts. Keep the accepted prefix; discard rejected cache.
- The accelerator needs host control, memory and links. A chip becomes a server.
- Servers need power and cooling. What changes when one model spans devices?
- Activations cross the fabric. Design the whole path for useful, timely tokens.
Try it before moving on
What extra communication does a model split across devices require?
Check your reasoning
The participating devices must exchange the activations needed to continue the model computation. Communication time joins compute and memory access in the end-to-end token latency budget.
Request laboratory
Before a token, there is a prompt.
Follow the same request through a shared pool, a prefill/decode split, and an optional draft-and-verify path. Change one assumption and find the point where an optimization stops paying for itself.
Synthetic examples with assumed timings. Keep the model, tokenizer and output-quality policy fixed. A length change does not predict prefill time: enter a matching measurement or an explicit estimate.
01 · Queue
20ms
Wait for admission
02 · Prefill
120ms
Process the prompt and first token
03 · KV handoff
0ms
KV stays with the request
04 · Decode
3,060ms
255 rounds after the first token
3.2 s to finish · 1× baseline speed
First token: 140 ms. Mean interval after it: 12 ms. Shared-pool ordinary decode baseline: 3.2 s.
Inspect the model and plan the next experiment
TTFT = initial queue + prefill + optional KV transfer + handoff wait. Transfer = prompt positions × KV bytes per position ÷ sustained one-way link rate. Completion = TTFT + rounds × time per round. This model holds the sampled first token until handoff finishes; emitting it earlier would move that wait to the interval before token two.
Ordinary rounds = output length − 1. Speculative rounds = ceiling((output length − 1) ÷ (mean accepted drafts + 1)); each costs draft length × draft-step time + target verification + commit/rollback. The final round may do excess work. Mean acceptance and rounded rounds are approximations, not a distribution or tail-latency simulation.
No transfer/compute overlap, prefix reuse, per-layer cache replication, EOS stopping, arrival process, pool capacity or contention model is included. Compare cold and warm prefixes separately, then replay concurrency 1, 8 and 32 in your project. Track exact KV ownership through cancellation and failed handoff.
Build a queue, cache controller, matrix tile or link endpoint that changes one measured term. Save the trace, numerical reference, implementation, tests and before/after timing. Carry its traffic and power into the chip-to-cluster worksheet below.
An inference system you can reason about
Follow one token. Design one part.
A token is the start of an engineering budget: bytes moved, arithmetic performed, memory occupied, and power delivered. Change the workload below and watch that budget travel from a chip to a cluster.
One more token
Decode appends one token to each sequence in a batch. All sequences share the weights; each has its own attention cache. This is decode, not prompt processing.
Evidence you can build
Save a workload contract: model shape, precision, context, batch size, and latency target.
Change the engineering budget
Illustrative dense decoder. One complete model per replica.
Model shape and sustained performance assumptions
Match these to your workload. Cache precision is independent of weight precision. Query heads must be a multiple of cache heads. Sustained percentages are your assumptions to replace with measurements.
Rack and cluster assumptions
Memory traffic sets this upper bound.
This is a calculated bound under your assumptions. Validate it with a measured kernel and end-to-end serving test.
Traffic per output token
20.06
GB / token
Matrix work per output token
80.237
GMAC / token
Resident model + batch
160.48
GB / 576 GB usable
Per-replica throughput bound
1,336
output tokens / s across the batch
Show every term in the calculation
- Weights: parameters × bits ÷ 8
- 139 GB shared across the batch; 17.375 GB read per output token after amortizing over 8 sequences.
- Cache bytes per position: 2 × layers × cache heads × head width × cache bits ÷ 8
- 320 KiB per sequence position. Read the 8192-position cache (2.6844 GB) and write one new position per output token.
- MACs per output token ≈ parameters + 2 × layers × context × query heads × head width
- Dense projections plus the QK and AV attention products. One MAC is one multiply-accumulate, conventionally two arithmetic operations. Compare throughput only at matching precision.
- Throughput ≤ min(sustained bandwidth ÷ bytes, sustained MAC/s ÷ MACs)
- Memory: 1,336 tokens/s. Arithmetic: 49,303.7 tokens/s. Capacity reserves 10% for other memory needs.
Carry it into a rack and cluster
Replicas are independent. Include each replica’s share of host and networking in its power input. PUE adds facility overhead; the result is a modeled steady-state load, not a thermal or electrical signoff.
Rack throughput bound
10,688.1
output tokens / s
Rack facility power
53.76
44.8 kW IT × 1.2 PUE
Cluster throughput bound
85,504.8
output tokens / s
Cluster facility power
430.08
kW
Facility energy at the throughput bound: 5.03 joules per output token. Reduced throughput increases energy per token if power stays fixed.
Take an experiment into your project
Paste this into your chapter notes or repository. Replace the assumptions with measurements as you build.
INFERENCE COMPONENT EXPERIMENT — illustrative decode estimate Contract: 69.5B dense parameters; 16-bit weights; batch 8; context 8192. Shape: 80 layers; 64 query / 8 cache heads; width 128; 16-bit cache. Traffic: 20.0597 GB/output token. Matrix arithmetic: 80.2374 GMAC/output token. Replica: 640 GB; 26800 GB/s at 100%; 3956 TMAC/s at 100%; 5600 W. Result: memory limit; 1,336 output tokens/s per replica; 430.1 kW modeled cluster facility power (8 replicas/rack × 8 racks, PUE 1.2). My component: [name and boundary]. My claim: [measurable improvement]. Evidence: workload.json, implementation, reproducible tests, measured traffic/latency/power, and a checksummed dataset manifest. Limits: no prefill, model sharding, interconnect contention, scheduling, or non-matrix arithmetic; 10% memory reserve. These are bounds, not benchmark results.
Units: GB = 10⁹ bytes; KiB = 1,024 bytes; TMAC/s = 10¹² multiply-accumulates per second. This model excludes prompt processing, sparse routing, non-matrix operations, quantization metadata, communication, and scheduling overhead. It assumes one weight stream per batch and one cache stream per sequence. It is a transparent worksheet to challenge with evidence.
Chapters
1. Turn request traces into a hardware mission
What does the workload physically demand, and how is a claim about silicon proved?
Two request traces become numbers a chip has to meet: bytes per token, first-token latency, tokens per second at a given concurrency. You learn how a chip is made, take a small design through the whole flow once, and learn to read the evidence a run leaves behind, so every later claim in the training is a file you can open.
- 1. What an Inference Part Has to MoveOn the first pass, define your two-request workload contract and component boundary. After the first-chip, RTL and verification chapters, return to derive a complete inference-part specification - HBM stack count and generation, memory beachfront in millimetres, interposer area, package power and core supply current - from the bytes one decode step moves, and identify which of bandwidth, capacity, beachfront or thermals binds each budget.~3 hours
- 2. How Chips Get MadeExplain, in an interview, who does what between an idea and packaged silicon: design, verification, physical design, signoff, foundry — and define PDK, tapeout, MPW, netlist, signoff, and PPA without notes.~1.5 hours
- 3. Your First Chip in One SittingTake a small design from RTL to GDSII, interpret the evidence at each stage, and identify how the same flow will implement the datapath or control block in your inference project.~2.5 hours
- 4. The Cockpit and the Chain of EvidenceNavigate every pane of a task, and distinguish passed, failed, tool-missing, and skipped by what each proves.~2 hours
- 5. Working With AI: The Engineer of RecordBrief an agent precisely, review AI-produced RTL and fixes against measured evidence, catch an unverified claim, and articulate where AI fails — a named career skill, not a footnote.~3 hours
2. Build the compute and KV-state path
Can the component compute the right numbers, at rate, under every stall and boundary?
The compute and KV-state path is built from the arithmetic up: number formats, synthesizable RTL, benches that find bugs, formal proof, the matrix engine, the processor beside it and the memory system that feeds them. Each chapter ends with a verified piece of your own component.
- 6. Number Formats and Quantization in SiliconSpecify and defend the numeric format set for an inference datapath — element format, scale granularity and accumulator width — from area you measured yourself rather than from a vendor's bit count.~3 hours
- 7. RTL That SynthesizesWrite and review synthesizable SystemVerilog — FSMs, FIFOs, arithmetic, CDC-safe resets — and direct an AI agent to write it while you stay the reviewer.~3 hours
- 8. Verification with SystemVerilog and UVMWrite self-checking SystemVerilog testbenches and understand the UVM architecture - agents, drivers, monitors, scoreboards, sequences - well enough to read, extend, and defend a UVM environment in an interview, and run an independent functional regression against your own RTL.~3.3 hours
- 9. Verification as a DisciplinePractise coverage thinking, regression discipline, and requirements tracing — the habits DV interviews actually test — across the flows you have already run.~3 hours
- 10. Formal Verification and Design for TestState what a bounded proof does and does not prove, and explain scan chains and ATPG well enough for a DFT interview question.~2.5 hours
- 11. Systolic Arrays and Matrix EnginesCompute what a fixed matrix array will actually achieve on a real transformer layer - row fill, column fill and fill-and-drain, each factor separately - then build a weight-stationary systolic tile that measures its own utilization in hardware and defend the number it reports.~3 hours
- 12. RISC-V and Custom AcceleratorsRead and extend a RISC-V-style datapath, and design a small domain-specific accelerator - the architecture pattern behind the industry's move from general-purpose processors to custom AI silicon, and the design vehicle your capstone builds on.~3.3 hours
- 13. Memory Systems: Where Inference Actually LivesSize the memory system for a stated inference serving target — stack count, on-die SRAM, KV placement and ECC policy — and defend the SRAM/HBM split from bandwidth and capacity budgets you derived and measured yourself.~3 hours
3. Prove the component at a physical operating point
What rate, area and power does the component really achieve, and what does signoff certify?
Synthesis, timing, clock domains, power, physical design and signoff replace the rate you assumed with the rate the silicon achieves at a real operating point, and teach exactly what a green signoff proves and what it does not.
- 14. Synthesis and Static TimingRead a synthesis report and a timing report cold, explain setup/hold/slack/critical path, and fix a setup violation.~3 hours
- 15. Clock Domains and CDCDesign safe clock-domain crossings - single-signal synchronizers, handshakes, and asynchronous FIFOs - and explain metastability and CDC verification the way interviewers demand, because CDC questions appear in almost every digital design interview.~2.5 hours
- 16. Low-Power SoC DesignApply the low-power toolkit - clock gating, power domains, UPF power intent, DVFS, and power-delivery reasoning - and read power reports critically, because power efficiency is a hard constraint in every AI data-center and edge deployment.~3 hours
- 17. PPA: The Trade-off LoopRun area-, timing-, and power-targeted iterations, compare metrics quantitatively, and defend a PPA decision the way a lead would ask you to.~2.5 hours
- 18. Floorplan to Routed SiliconExplain floorplanning, placement, CTS, and routing as distinct problems, and read openroad_metrics.json and a routed DEF like a PD engineer.~3.3 hours
- 19. Signoff: DRC, LVS, and Physical VerificationDebug a DRC violation from the deck to the DRM rule, explain what 'match uniquely' means, and state exactly what a SIGNOFF-CLEAN verdict certifies.~3 hours
4. Supply the chip with data, power and cooling
Can the part receive its bytes and current and reject its heat under bursty prefill and sustained decode?
A digital part lives at an analog edge. High-speed links, an amplifier designed from a specification, extraction, mixed-signal integration, the SerDes lane and the power and thermal budget decide whether the chip can be fed, powered and cooled at the rates the workload demands.
- 20. High-Speed Interfaces and SerDesUnderstand how data actually enters and leaves a chip - parallel buses to multi-gigabit SerDes, the PCIe/CXL/DDR families, and the signal-integrity reasons the physical layer dominates AI-cluster performance - and take an interface block through timing closure.~3 hours
- 21. Analog: Schematic Thinking and SimulationRead a SPICE netlist fluently, run DC/AC/transient analyses against a spec, and explain MOSFET operating regions and biasing from your own simulation data.~3.3 hours
- 22. The Op-Amp From a Spec SheetTake a constraining op-amp spec — gain, GBW, phase margin, power — to simulated closure, and walk the noise and stability budget aloud, interview-style.~4 hours
- 23. Analog Layout, Matching, and ParasiticsExplain matching, common-centroid, and parasitic effects, and carry a layout through DRC, LVS, and post-extraction re-simulation.~3.7 hours
- 24. Mixed-Signal IntegrationIntegrate analog and digital in one design with an explicit boundary contract, and run AMS verification across the domain wall.~3.3 hours
- 25. What Is Analog on a Digital ChipTake apart the analog edge of a part everyone calls digital - the SerDes lane as a circuit, the power delivery network, the sensors and references behind it - say for every block whether it is synthesized from RTL or drawn by hand in a schematic editor, and defend from your own simulated and computed numbers why the hand-drawn half does not shrink at the next node.~3 hours
- 26. Power, Thermal, and Delivery at a KilowattBuild a power budget for a kilowatt-class inference part, name which constraint actually binds it — delivery, junction temperature, or the rack feed — and defend the operating point you chose from your own arithmetic.~2.5 hours
5. Serve the workload across hosts, racks and clusters
What limits useful tokens per joule and per dollar once many parts, hosts and links serve many requests?
Chiplets, interconnect, the rack, the network, the software stack and bring-up turn one component into a serving system, and benchmarking turns that system into a cost per token you can defend.
- 27. Interconnect, Chiplets, and Making One Model Span Many DiesPartition a stated model across a stated topology, compute the per-token collective traffic and the latency each collective adds to a decode step, and defend where the die, package and network boundaries should fall.~3 hours
- 28. The Rack Is the ComputerSize the unit that actually serves a model - the rack - from its three-phase feed and 48 V busbar through its cooling loop, its two fabrics, its optics and its PCIe and CXL lanes, and defend where a serving system's memory lives across HBM, DDR, CXL and SSD.~3.2 hours
- 29. The Hardware-Software ContractSpecify, in gates, the command and debug interface a matrix engine has to present — descriptor ring, doorbell, completion path, ordering and coherence rules, address generator, semaphores and telemetry — and defend every clause of it with your own measured area and slack numbers, including where you drew the fixed-function line and what you kept that delivers no tokens per second at all.~3 hours
- 30. Hardware-Software Co-Design and Bring-UpWork the seam between silicon and software: C/C++/Python for pre-silicon modeling and post-silicon bring-up, what FPGA prototyping and emulation are for, and how a chip is actually brought to life on a bench.~2.5 hours
- 31. Benchmarking Inference, and the Economics That Decide the DesignCompute cost per million tokens for a candidate inference part at a stated SLO, context length and number format, name which input of your own cost model dominates the answer, and defend a build-or-buy recommendation with a break-even volume you derived rather than quoted.~3 hours
6. Release and defend the complete inference design
Can another engineer reproduce every claim, and what must be published for that to be true?
The accelerator is assembled end to end, the economics of tapeout are counted, the work is made employable, and the capstone is built, published and credentialed so that every number in your dossier is a file someone else can open.
- 32. Build It: An Attention Engine Through the Full FlowTake a tiled INT8 attention datapath - weight-stationary array, double-buffered weight memory, accumulator memory, requantize stage and a fully specified register map - from a pasted spec through a 4,571-check regression and every flow stage to signoff, then defend a tokens-per-second figure derived from your own measured cycle count.~3 hours
- 33. Tapeout Mechanics, Bring-up, and Chip EconomicsExplain what happens between GDS and a working part — shuttles, packaging, bring-up — and reason about ASIC cost and schedule like someone who has shipped.~2.5 hours
- 34. Employability: Interviews, Portfolio, and ProofMap your artifacts to job descriptions, rehearse the questions each track's interviewers ask, and make your competence publicly checkable on this platform.~2.5 hours
- 35. Capstone I: Design and Sign Off Your ChipPropose, design, and carry your own chip through a complete flow to a clean signoff verdict.~11 hours
- 36. Capstone II: Publish, Credential, and the FlywheelTurn your capstone into public, citable proof: a trajectory dataset, a chip space, and a personalized model recipe template under your own name — then claim the server-verified credential.~3 hours
Your Inference Component and Chip-to-Cluster Portfolio
Build a bounded compute, attention, memory, interconnect, telemetry or system-control component for your request traces. Deliver the implementation and a chip-to-cluster dossier that follows prefill, KV state and decode through host software, package interfaces, rack power and cluster placement. Compare alternatives at matched request distribution, concurrency, cache locality, quality and TTFT/TPOT; include useful tokens, rejected work, facility energy and cost. Carry your selected inference component through specification, independent verification, synthesis, PPA iteration, physical design and DRC/LVS/PV signoff. The MAC-array chapter supplies a complete reference design, including a memory-mapped interface and multi-clock integration; use it to study those requirements without abandoning your chosen component. Analog and mixed-signal companion projects develop the electrical interface and supply optional flow endorsements. Release source, measured reports and the remaining limitations together, so another engineer can reproduce and question each claim.
- A completed silicon run per chosen flow with GDSII, TAPEOUT_REPORT.md, DESIGN_RECORD.{json,md}, timing/DRC/LVS/PV reports, and the FIRST_SILICON handoff documents, plus the checksummed handoff bundle
- Signed stage receipts on every review/handoff stage, naming the student
- A documented PPA comparison across at least two objective-targeted iterations
- HUB DATASET <username>/<project>-trajectories (public): the verified run evidence (signoff files and stage reports from the final clean run; iteration history stays in the task branch) — published by the platform from the run itself, its card naming the flow, the sha256 digest of the exact files published, and a run reference shared by all three of your artifacts so a reader can tell they came from one run
- HUB SPACE <username>/<project>-chip (public): captured GDS, DEF, LEF and available netlists with a design card. These are browsable physical-design files for inspection in a layout tool; this publication does not execute RTL in the browser.
- HUB MODEL <username>-chipcraft-v1 (public): a personalized fine-tune recipe template — recipe.json identifying the student's verified-run corpus and proposed method; no trained weights, with a compatible base model, example preparation, training configuration and evaluation plan still to be supplied
- The profile credential on your public profile page and your Talent Network profile, stating the flows completed and linking the three public artifacts
- Versioned request traces and an architecture contract covering model shape, precision, prefill, KV ownership, decode and optional draft/verify acceptance and rollback, with a bit-accurate reference, interface specification and requirement-to-test matrix.
- A chip/package/board/host/rack/cluster dossier with software integration, bandwidth, power, thermal, reliability, cache locality and capacity budgets. Compare quality and TTFT/TPOT at the same request distribution and concurrency, including useful output tokens per facility energy and cost; label measured component results and system projections.
- A reproducible dataset release manifest with provenance, licenses, file sizes and hashes, download verification and explicit limitations. Identify the process and checks covered by signoff evidence and the qualification, fabrication and manufacturing tests still required for deployment.
How completion is verified: Every chapter ends with a check for understanding; passing it is what marks that chapter complete. Browsing remains open in any order. Check retries are free and unlimited and ask only the missed questions. The credential additionally requires every chapter complete, the required synchronous run and each claimed companion flow signed off clean, and three qualifying capstone artifacts still accessible under the publication rules and tied to those runs. A completed flow with blocked physical verification checks is insufficient. Publish or republish from the final clean run before requesting the credential. The cited artifacts are re-checked when the profile is displayed, so later visibility or provenance changes can remove an unsupported credential.
