The hub for chip design: datasets, evals, and the models trained on them

Every other field got one place where the data, the benchmarks that grade it, and the models trained on both live together. Silicon never did. ActGen is that place — and unlike a passive archive, the data is generated here, by experts correcting AI inside a real chip-design workflow.

The platform runs an AI-driven EDA flow where every sign-off gate is scored by rubrics and evals, and where experts working with AI assistance ultimately correct its output. Those corrections, trajectories, and verdicts become training data: irreversibly scrubbed of secrets at capture, then consent-gated and de-identified before anything is exported, and machine-verified by open EDA tooling. The evals that graded the work ship beside it, so the thing you train on and the thing that judges it come from one place.

Everything here is organized by role in the design flow rather than by model modality — pre-silicon verification, post-silicon debug, EDA and APR implementation, orchestration, triage, security — because that is the axis a rubric can attach to. Browse the roles.

Partial sponsorship: the industry-wide Open Layout Benchmark
We work with frontier labs on frontier data, rubrics, and evals for chip-design tasks. Now we are building a public dataset for GDSII and the data around it, playing the role a "layout CVDP" would — the layout counterpart of CVDP, the well-known benchmark of Verilog design problems — capturing real layout tasks from the companies that live them.This is a multi-company collaborative effort, and each participant contributes as little or as much as it wants. Your tier decides how far the dataset is tailored to your tasks and the features you want in it, and how much sponsor-only supplementary data you receive beyond the public release. Sponsoring organizations are recognized publicly — like a conference sponsorship — once they confirm their listing.What goes in: paired task-and-outcome data across the layout stack — DRC-violation repair episodes with full tool traces and the final clean run, LVS mismatch diagnosis, congestion-driven reroutes, timing-driven ECO implementation, floorplan and placement optimization with before-and-after QoR, and expert corrections of AI-produced cell layouts against golden GDSII — each row carrying its DEF/LEF context, the rubric that graded it, and a machine-verified outcome from open EDA tooling.Sponsor-only data is delivered on the platform itself: your supplement — the held-out partition the public release never contains — is published as a gated hub dataset and your account is granted access directly, in the same operator approval that releases it. Your organization can also maintain fully private datasets on the hub: private repositories are visible to their owner alone, and nothing you keep private is published without the operator's explicit final approval.Why sponsor? A public benchmark of your hardest layout tasks referees internal and external agents against work that actually matters to you, and it points the research community's weight at exactly those problems. Bring us the tasks today's AI cannot solve, and track your own progress objectively: what gets measured gets better.

Supporter

$10,000

Back the public release. Your organization is credited among the sponsors when you confirm your listing.

Task Sponsor

$50,000

Bring your own layout tasks: the dataset captures task families and features meaningful to you, and you receive sponsor-only supplementary data beyond the public release.

Anchor Sponsor

$150,000

Shape the benchmark: first call on task selection and dataset features, the largest sponsor-only data grant, and lead recognition across the public release.

Tiers are starting points — every amount between the published bounds works, and the form below is yours to edit before anything is charged.

The referee kit — rubrics & evals for agent fleets

A versioned pack of sign-off gate rubrics and qualitative intent criteria with live judge calibration: written success definitions, machine-measured thresholds, scoring semantics, and an override protocol that makes verdicts comparable across teams and agents. The rubrics learn — expert disagreements drive LLM-drafted, human-approved revisions, and calibration tracks whether agreement improves.

It is runnable, not just readable. The pack carries the same scoring constants and algorithm this platform grades itself with, plus a worked example to check your implementation against, so your team can score their own agents and get the numbers we would get. Every revision ships with a changelog — what the rubric said before, what it says now, and how each wording performed — and every agreement figure carries the counts behind it and states which set it was measured over, because a rate without its denominator is not evidence. Refereed runs the operator records are published on the public leaderboard, each score pinned to the pack version it measured.

And it is yours, not a newsletter. Your rubric is scoped to you: you can reword a criterion or add one we do not ship (power-intent coverage, CDC closure, DFT insertion), and nothing you negotiate touches the pack any other party holds. Where you use our wording you get our calibration with it; where you use your own, your pack says so and omits the figures — because a rate we measured under different wording, printed beside yours, would be a true number making a false claim.

The data catalog — accrued, consent-gated packages

Task conversations, agentic SFT turns, preference pairs, repair triples, expert corrections paired with the AI output they overrode, whole RL trajectories with machine-verified terminal rewards, and a simulator-verified RTL corpus — every user-derived row consent-gated with retroactive revocation, pseudonymized, and secret-scrubbed, delivered with machine-generated datasheets down to per-file hashes.

Training and evaluation are separate partitions, disjoint by construction rather than by convention — a slice never carries held-out benchmark rows unless you ordered the held-out set. Every datasheet states which partition you hold, what it licenses, and a contamination attestation: verbatim duplication inside the package, measured overlap against each held-out corpus with the counts behind it, and an explicit list of the public suites we did not check, so you know exactly which diligence is still yours.

Data on demand — commissioned expert production

Order a specific slice — a package, a row volume, a coverage brief — and experts produce it through normal platform usage on a time-and-materials basis, with server-measured fulfillment, delivery snapshots, and machine-generated invoices.

Sponsor a dataset directly

Know what you need? Fund a commissioned dataset now — the order lands with the operator, who staffs it from the consent-gated expert roster and confirms scope with you before production starts.

Sponsor a dataset
Fund a commissioned dataset directly — describe what you need, choose a budget between $5,000 and $250,000, and pay here. Requires a signed-in account.

Start on the platform

Everything here is self-serve: create an account to browse the catalogs, run the evals against your own agents, and sponsor or license directly — samples and datasheets live with each offering.

ActGen — the hub for chip-design datasets, evals, and the models trained on them