ChipGPT Engineering RLE

Know whether an AI agent can engineer silicon—not merely write Verilog.

A private, executable environment for evaluating and improving agents across RTL, design verification, firmware, RTOS, security, formal, and coverage engineering.

  • 70 task candidates
  • 70 distinct concept groups
  • Golden / broken / oracle validated
  • Private executable grading

Oracle validation proves task integrity; full multi-model calibration is pending.

The gap

Writing Verilog is not the same as engineering silicon.

Prompt / response HDL benchmarks score a single generated answer. Real engineering work is multi-step, tool-driven, and judged by behavior the author cannot see. These are the things a one-shot benchmark cannot measure:

Repository navigation

Real work starts by finding the right files in a codebase, not by reading a self-contained prompt.

Iterative tool use

Engineers compile, simulate, read failures, and try again. A single response measures none of that.

Hidden behavior

Visible checks can be satisfied by pattern-matching. Protected tests check the behavior that actually matters.

Regressions

A fix that breaks adjacent behavior is not a fix. Non-regression has to be graded, not assumed.

Test tampering

An agent that weakens a test, adds a waiver, or hard-codes a fixture can look like it passed.

Repeatability and cost

One lucky success is not a capability. Repeated runs, tokens, time, and compute all belong in the result.

Engineering RLE is the concrete evaluation layer of ChipGPT’s Reliability stack: it measures and improves AI co-workers on executable engineering tasks.

Portfolio

70 task candidates across five engineering lanes.

Each task supplies an isolated workspace, a precise engineering objective, bounded tools, visible development checks, and protected executable grading.

Engineering RLE task portfolio by lane
Engineering laneTasks
RTL semantic repair20
Security and safety implementation10
Firmware and RTOS10
Mutation-scored DV generation10
Coverage engineering and evidence20
Total70

How tasks are graded

An executable evaluation flow, not a subjective review.

A task passes only when every trusted correctness gate passes: integrity, build, private target behavior, architecture / mutation, non-regression, deterministic replay, and scope compliance. A weighted partial reward cannot override a failed hard gate.

  1. 01

    Pinned source + specification

    Fixed revision, precise engineering objective

  2. 02

    Isolated agent workspace

    Fresh, resource-bounded, with bounded tools

  3. 03

    RTL / DV / software patch

    Agent edits RTL, DV, C, assembly, or properties

  4. 04

    Protected grader

    Integrity → build → hidden behavior → non-regression → replay

  5. 05

    Report

    Capability · reliability · quality · efficiency

Network-disabled OCI

The production backend is designed for digest-pinned, network-disabled OCI runs. Only campaigns actually run there qualify as production evidence.

Credential separation

Model-provider and grader credentials are kept separate from the agent-visible workspace.

Protected tests

Private tests, mutants, and oracle repairs stay outside the agent workspace. Grader or test tampering is rejected before private grading.

Deterministic replay

A passing submission must replay to the same verdict. Replay is a hard gate, not a bonus.

Bounded tools

Agents get read, search, patch, and shell operations within declared allowed paths and resource limits.

Synthesis, place-and-route, physical PPA, scan insertion, and ATPG are outside this tranche. Static synthesizability and lint checks may still run.

What you learn

Five dimensions of an agent, measured separately.

Correctness dominates efficiency: a cheaper failure is still a failure.

Capability

Family-macro pass@1

Does the agent solve the task on its first attempt, macro-averaged across task families?

Reliability

Repeated-run success

Does it succeed again, or only once by chance?

Verification strength

Hidden mutant detection and false-failure control

Do the agent's tests catch real bugs without failing on unrelated changes?

RTL quality

Lint, latches, X behavior, interfaces, maintainability

Is the result something an engineer would accept into the tree?

Efficiency

Tokens, time, tools, simulation compute, retries, patch size

What did the result cost to produce?

Evidence

What has been validated so far.

These results validate task construction and the grading path. They are not model scores.

  • 70/70 golden states pass candidate validation.
  • 70/70 deliberately broken states are detected.
  • 70/70 oracle states pass.
  • Three complete validations produced the same normalized verdict digest.
  • Oracle baselines cover all 70 tasks.
  • Granite 4.2 8B solved one ORI development smoke episode.One successful development smoke task—not a 70-task score or a publishable model comparison.

Design-partner pilot

Start with five tasks and one agent configuration.

Pilot includes

  • Five representative tasks
  • One model / agent configuration
  • Three attempts per task
  • Standard integration and hosted private grading
  • Normalized results and a capability report
  • Two technical reviews

Design-partner pilot anchor

Five-task design-partner pilots start at $50,000

Subject to final scope, license approval, and signed agreement.

Your model can stay where it is.At the starting price, the customer’s model may remain behind its own secured endpoint while task execution and private grading stay ChipGPT-managed.

Other deployments are scoped separately. Customer-VPC and connected on-premise execution are separately scoped. Fully air-gapped execution is a custom option subject to technical, licensing, and security readiness approval.

Credit toward the full suite. If you expand to a complete-suite agreement within the contractually stated window, the task-license portion of the pilot is credited. Deployment fees are non-creditable.

Rights.The standard pilot includes a 90-day license for internal evaluation and internal model training / improvement. You retain your generated weights, outputs, and patches, subject to the agreement’s underlying rights. Continued task access, publication, and redistribution require a separate written grant.

API / compute charges, commercial EDA tools, and custom integrations are separate unless stated in the order form.

FAQ

Questions buyers ask.

What does 70/70 mean?

All 70 task candidates pass golden / broken / oracle construction validation: every golden state passes its candidate checks, every deliberately broken state is detected, and every oracle repair passes. Three complete validation runs produced the same normalized verdict digest.

This proves executable task construction, not model difficulty. 70/70 is not a model score.

Did Granite pass all tasks?

No. No portfolio score exists. Granite 4.2 8B solved one ORI task in one successful development smoke episode. Other attempts on that task were not all successful, and Granite has not been run across all 70 tasks.

Does this run synthesis or report PPA?

No. This tranche excludes synthesis, place-and-route, gate count, physical timing, power, scan insertion, and ATPG. RTL structure may be reported as a review diagnostic, not a PPA reward.

How are private graders protected?

Allowed and protected paths are declared per task. Grader or test tampering, prohibited paths, oversized patches, new waivers, force/release bypasses, and other integrity violations are rejected before private grading.

Under the standard hosted offer, protected tests, mutants, and oracle repairs remain behind the grader boundary; buyers receive normalized results and permitted evidence.

Can model weights stay in our environment?

The architecture supports controlled evaluation patterns where customer models remain behind the customer's endpoint. The exact deployment, logging, data retention, and grader boundary are agreed during security design.

Customer-VPC and connected on-premise execution are separately scoped and are not included in the pilot by default.

Can tasks be used for training?

Yes. The standard pilot grants one named customer organization or business unit a 90-day right to use its five licensed tasks for internal evaluation and internal model training / improvement, subject to applicable source and provider terms.

The customer retains its generated model weights, outputs, and patches, subject to underlying task, environment, and third-party rights. Continued task access, redistribution, and publication are not included. Training on the suite is not a guarantee of model improvement.

What is included in the pilot?

Five tasks, one model / agent configuration, three attempts per task, standard integration, hosted private grading, normalized evidence, one capability report, and two technical reviews.

API / compute, commercial EDA tools, custom integrations, deployment changes, and expanded rights are separate unless stated in the order form.

Do you use an LLM judge?

An LLM review may be advisory, but it is never a correctness hard gate. The authoritative result comes from executable checks and explicit policies.

Put your model in front of executable silicon-engineering work.

Request a 25-minute technical briefing about a paid five-task design-partner pilot. We’ll walk through task design, grading, and how your model or agent harness would connect.

Request an RLE Briefing

Deployment choices are for discovery only; listing an option does not mean it is available for every engagement. We use these details only to respond to your request. Please don't include confidential source, model weights, or credentials.

or email connect@chipgpt.ai