ChipGPT Engineering RLE
Know whether an AI agent can engineer silicon—not merely write Verilog.
A private, executable environment for evaluating and improving agents across RTL, design verification, firmware, RTOS, security, formal, and coverage engineering.
- 70 task candidates
- 70 distinct concept groups
- Golden / broken / oracle validated
- Private executable grading
Oracle validation proves task integrity; full multi-model calibration is pending.
The gap
Writing Verilog is not the same as engineering silicon.
Prompt / response HDL benchmarks score a single generated answer. Real engineering work is multi-step, tool-driven, and judged by behavior the author cannot see. These are the things a one-shot benchmark cannot measure:
Repository navigation
Real work starts by finding the right files in a codebase, not by reading a self-contained prompt.
Iterative tool use
Engineers compile, simulate, read failures, and try again. A single response measures none of that.
Hidden behavior
Visible checks can be satisfied by pattern-matching. Protected tests check the behavior that actually matters.
Regressions
A fix that breaks adjacent behavior is not a fix. Non-regression has to be graded, not assumed.
Test tampering
An agent that weakens a test, adds a waiver, or hard-codes a fixture can look like it passed.
Repeatability and cost
One lucky success is not a capability. Repeated runs, tokens, time, and compute all belong in the result.
Engineering RLE is the concrete evaluation layer of ChipGPT’s Reliability stack: it measures and improves AI co-workers on executable engineering tasks.
Portfolio
70 task candidates across five engineering lanes.
Each task supplies an isolated workspace, a precise engineering objective, bounded tools, visible development checks, and protected executable grading.
| Engineering lane | Tasks |
|---|---|
| RTL semantic repair | 20 |
| Security and safety implementation | 10 |
| Firmware and RTOS | 10 |
| Mutation-scored DV generation | 10 |
| Coverage engineering and evidence | 20 |
| Total | 70 |
How tasks are graded
An executable evaluation flow, not a subjective review.
A task passes only when every trusted correctness gate passes: integrity, build, private target behavior, architecture / mutation, non-regression, deterministic replay, and scope compliance. A weighted partial reward cannot override a failed hard gate.
- 01
Pinned source + specification
Fixed revision, precise engineering objective
- 02
Isolated agent workspace
Fresh, resource-bounded, with bounded tools
- 03
RTL / DV / software patch
Agent edits RTL, DV, C, assembly, or properties
- 04
Protected grader
Integrity → build → hidden behavior → non-regression → replay
- 05
Report
Capability · reliability · quality · efficiency
Network-disabled OCI
The production backend is designed for digest-pinned, network-disabled OCI runs. Only campaigns actually run there qualify as production evidence.
Credential separation
Model-provider and grader credentials are kept separate from the agent-visible workspace.
Protected tests
Private tests, mutants, and oracle repairs stay outside the agent workspace. Grader or test tampering is rejected before private grading.
Deterministic replay
A passing submission must replay to the same verdict. Replay is a hard gate, not a bonus.
Bounded tools
Agents get read, search, patch, and shell operations within declared allowed paths and resource limits.
Synthesis, place-and-route, physical PPA, scan insertion, and ATPG are outside this tranche. Static synthesizability and lint checks may still run.
What you learn
Five dimensions of an agent, measured separately.
Correctness dominates efficiency: a cheaper failure is still a failure.
Capability
Family-macro pass@1
Does the agent solve the task on its first attempt, macro-averaged across task families?
Reliability
Repeated-run success
Does it succeed again, or only once by chance?
Verification strength
Hidden mutant detection and false-failure control
Do the agent's tests catch real bugs without failing on unrelated changes?
RTL quality
Lint, latches, X behavior, interfaces, maintainability
Is the result something an engineer would accept into the tree?
Efficiency
Tokens, time, tools, simulation compute, retries, patch size
What did the result cost to produce?
Evidence
What has been validated so far.
These results validate task construction and the grading path. They are not model scores.
- 70/70 golden states pass candidate validation.
- 70/70 deliberately broken states are detected.
- 70/70 oracle states pass.
- Three complete validations produced the same normalized verdict digest.
- Oracle baselines cover all 70 tasks.
- Granite 4.2 8B solved one ORI development smoke episode.One successful development smoke task—not a 70-task score or a publishable model comparison.
Design-partner pilot
Start with five tasks and one agent configuration.
Pilot includes
- Five representative tasks
- One model / agent configuration
- Three attempts per task
- Standard integration and hosted private grading
- Normalized results and a capability report
- Two technical reviews
Design-partner pilot anchor
Five-task design-partner pilots start at $50,000
Subject to final scope, license approval, and signed agreement.
Your model can stay where it is.At the starting price, the customer’s model may remain behind its own secured endpoint while task execution and private grading stay ChipGPT-managed.
Other deployments are scoped separately. Customer-VPC and connected on-premise execution are separately scoped. Fully air-gapped execution is a custom option subject to technical, licensing, and security readiness approval.
Credit toward the full suite. If you expand to a complete-suite agreement within the contractually stated window, the task-license portion of the pilot is credited. Deployment fees are non-creditable.
Rights.The standard pilot includes a 90-day license for internal evaluation and internal model training / improvement. You retain your generated weights, outputs, and patches, subject to the agreement’s underlying rights. Continued task access, publication, and redistribution require a separate written grant.
API / compute charges, commercial EDA tools, and custom integrations are separate unless stated in the order form.
FAQ
Questions buyers ask.
What does 70/70 mean?
All 70 task candidates pass golden / broken / oracle construction validation: every golden state passes its candidate checks, every deliberately broken state is detected, and every oracle repair passes. Three complete validation runs produced the same normalized verdict digest.
This proves executable task construction, not model difficulty. 70/70 is not a model score.
Did Granite pass all tasks?
No. No portfolio score exists. Granite 4.2 8B solved one ORI task in one successful development smoke episode. Other attempts on that task were not all successful, and Granite has not been run across all 70 tasks.
Does this run synthesis or report PPA?
No. This tranche excludes synthesis, place-and-route, gate count, physical timing, power, scan insertion, and ATPG. RTL structure may be reported as a review diagnostic, not a PPA reward.
How are private graders protected?
Allowed and protected paths are declared per task. Grader or test tampering, prohibited paths, oversized patches, new waivers, force/release bypasses, and other integrity violations are rejected before private grading.
Under the standard hosted offer, protected tests, mutants, and oracle repairs remain behind the grader boundary; buyers receive normalized results and permitted evidence.
Can model weights stay in our environment?
The architecture supports controlled evaluation patterns where customer models remain behind the customer's endpoint. The exact deployment, logging, data retention, and grader boundary are agreed during security design.
Customer-VPC and connected on-premise execution are separately scoped and are not included in the pilot by default.
Can tasks be used for training?
Yes. The standard pilot grants one named customer organization or business unit a 90-day right to use its five licensed tasks for internal evaluation and internal model training / improvement, subject to applicable source and provider terms.
The customer retains its generated model weights, outputs, and patches, subject to underlying task, environment, and third-party rights. Continued task access, redistribution, and publication are not included. Training on the suite is not a guarantee of model improvement.
What is included in the pilot?
Five tasks, one model / agent configuration, three attempts per task, standard integration, hosted private grading, normalized evidence, one capability report, and two technical reviews.
API / compute, commercial EDA tools, custom integrations, deployment changes, and expanded rights are separate unless stated in the order form.
Do you use an LLM judge?
An LLM review may be advisory, but it is never a correctness hard gate. The authoritative result comes from executable checks and explicit policies.
Put your model in front of executable silicon-engineering work.
Request a 25-minute technical briefing about a paid five-task design-partner pilot. We’ll walk through task design, grading, and how your model or agent harness would connect.