Repair loop result
The first structured, reproducible benchmark for AI-assisted hardware design at RTL. Structured error feedback — not raw logs — is what makes the repair loop work. FSM syntax validity: 0% → 66.7% with constrained prompts alone.
scroll to expand
This establishes the first structured, reproducible benchmark for AI-assisted hardware design at RTL (Register Transfer Level). LLMs are given natural-language circuit specifications and must produce synthesisable Verilog HDL — evaluated across 50 tasks spanning combinational logic, sequential circuits, FSMs, and complex mixed designs. The framework goes beyond syntax checking: a semantic repair loop feeds simulation failures back to the model with structured error context, enabling iterative self-correction.
Conducted across 12 benchmark iterations — from a 5-task initial test to the full 50-task scope — evolving the pipeline from basic few-shot prompting through constrained prompts, post-processing, and finally semantic-aware iterative refinement with waveform analysis and formal equivalence checking.
When the simulator rejects generated Verilog, the naive approach is to dump the raw error log back to the model. This project takes a different approach: a feedback generator extracts structured error context — identifying whether the failure was a syntax issue, a port spec mismatch, an FSM state encoding problem, or a functional simulation failure — and formats it as targeted repair instructions rather than raw compiler output.
The pipeline evolved across 12 benchmark iterations. Phase 2 introduced constrained prompts with exact module and port name specifications, which alone produced a FSM breakthrough: StarCoder2 went from 0% to 66.7% syntax validity on FSM tasks. Phase 4 added AST-based repair and waveform diffing. Phase 5 introduced category-aware prompt templates for FSM and mixed designs.
The dataset is intentionally capped at 50 curated tasks — 23 combinational, 14 sequential, 8 FSM, 5 mixed — prioritising depth and analysis quality over breadth. Docker container enables full reproducibility on any platform.
# feedback_generator.py — structured error extraction def generate_feedback(task, verilog, sim_result): errors = sim_result['errors'] category = task['category'] # Identify failure mode if 'syntax error' in errors: return build_syntax_feedback(errors, task) elif 'port mismatch' in errors: expected = task['inputs'] + task['outputs'] return build_port_feedback(errors, expected) elif category == 'fsm': # Category-aware: FSM state encoding issues return build_fsm_feedback( errors, sim_result['waveform_diff'] ) return build_functional_feedback( errors, sim_result['expected_vs_actual'] )
Syntax validity and functional correctness across 3 models after Phase 5 (strict mode, waveform + formal enabled, up to 6 repair iterations per task).
Evolution across 12 benchmarks — functional correctness improvement at key phases.
| Phase | Methodology | Pass Rate |
|---|---|---|
| Phase 1 | Few-shot prompting | ~0% |
| Phase 2 | Constrained prompts + post-processing | ~25% |
| Phase 4 | Semantic repair + waveform analysis | ~42% |
| Phase 5 | Category-aware templates + micro-repair | 51.8% |
Functional correctness on FSM tasks remains the hardest category — state encoding ambiguity is a fundamental LLM limitation for hardware description, not a prompt engineering problem. The paper discusses this as an open problem.