The hub for chip design: datasets, evals, and the models trained on them
Every other field got one place where the data, the benchmarks that grade it, and the models trained on both live together. Silicon never did. ActGen is that place — and unlike a passive archive, the data is generated here, by experts correcting AI inside a real chip-design workflow.
The platform runs an AI-driven EDA flow where every sign-off gate is scored by rubrics and evals, and where experts working with AI assistance ultimately correct its output. Those corrections, trajectories, and verdicts become training data: irreversibly scrubbed of secrets at capture, then consent-gated and de-identified before anything is exported, and machine-verified by open EDA tooling. The evals that graded the work ship beside it, so the thing you train on and the thing that judges it come from one place.
Everything here is organized by role in the design flow rather than by model modality — pre-silicon verification, post-silicon debug, EDA and APR implementation, orchestration, triage, security — because that is the axis a rubric can attach to. Browse the roles.
Supporter
$10,000
Back the public release. Your organization is credited among the sponsors when you confirm your listing.
Task Sponsor
$50,000
Bring your own layout tasks: the dataset captures task families and features meaningful to you, and you receive sponsor-only supplementary data beyond the public release.
Anchor Sponsor
$150,000
Shape the benchmark: first call on task selection and dataset features, the largest sponsor-only data grant, and lead recognition across the public release.
Tiers are starting points — every amount between the published bounds works, and the form below is yours to edit before anything is charged.
The referee kit — rubrics & evals for agent fleets
A versioned pack of sign-off gate rubrics and qualitative intent criteria with live judge calibration: written success definitions, machine-measured thresholds, scoring semantics, and an override protocol that makes verdicts comparable across teams and agents. The rubrics learn — expert disagreements drive LLM-drafted, human-approved revisions, and calibration tracks whether agreement improves.
It is runnable, not just readable. The pack carries the same scoring constants and algorithm this platform grades itself with, plus a worked example to check your implementation against, so your team can score their own agents and get the numbers we would get. Every revision ships with a changelog — what the rubric said before, what it says now, and how each wording performed — and every agreement figure carries the counts behind it and states which set it was measured over, because a rate without its denominator is not evidence. Refereed runs the operator records are published on the public leaderboard, each score pinned to the pack version it measured.
And it is yours, not a newsletter. Your rubric is scoped to you: you can reword a criterion or add one we do not ship (power-intent coverage, CDC closure, DFT insertion), and nothing you negotiate touches the pack any other party holds. Where you use our wording you get our calibration with it; where you use your own, your pack says so and omits the figures — because a rate we measured under different wording, printed beside yours, would be a true number making a false claim.
The data catalog — accrued, consent-gated packages
Task conversations, agentic SFT turns, preference pairs, repair triples, expert corrections paired with the AI output they overrode, whole RL trajectories with machine-verified terminal rewards, and a simulator-verified RTL corpus — every user-derived row consent-gated with retroactive revocation, pseudonymized, and secret-scrubbed, delivered with machine-generated datasheets down to per-file hashes.
Training and evaluation are separate partitions, disjoint by construction rather than by convention — a slice never carries held-out benchmark rows unless you ordered the held-out set. Every datasheet states which partition you hold, what it licenses, and a contamination attestation: verbatim duplication inside the package, measured overlap against each held-out corpus with the counts behind it, and an explicit list of the public suites we did not check, so you know exactly which diligence is still yours.
Data on demand — commissioned expert production
Order a specific slice — a package, a row volume, a coverage brief — and experts produce it through normal platform usage on a time-and-materials basis, with server-measured fulfillment, delivery snapshots, and machine-generated invoices.
Sponsor a dataset directly
Know what you need? Fund a commissioned dataset now — the order lands with the operator, who staffs it from the consent-gated expert roster and confirms scope with you before production starts.
Start on the platform
Everything here is self-serve: create an account to browse the catalogs, run the evals against your own agents, and sponsor or license directly — samples and datasheets live with each offering.
