Portfolio

Repair loop result

0% Phase 1 baseline
few-shot prompting
51.8% Phase 5 with semantic
repair loop · 1,610 runs
3 LLMs
50 tasks
12 benchmark
iterations

The first structured, reproducible benchmark for AI-assisted hardware design at RTL. Structured error feedback — not raw logs — is what makes the repair loop work. FSM syntax validity: 0% → 66.7% with constrained prompts alone.

scroll to expand

Python Ollama Verilog / HDL Verilator Icarus Verilog Docker LangGraph Paper Under Review
arXiv:2607.22759 → GitHub
1,610 Total benchmark runs across 12 iterations
0→51.8% Functional pass rate lifted by repair loop
50 Curated HDL tasks across 4 circuit categories
3 LLMs evaluated — 8B, 7B, and 1.1B parameters
What this is

AI writing hardware — measured rigorously

This establishes the first structured, reproducible benchmark for AI-assisted hardware design at RTL (Register Transfer Level). LLMs are given natural-language circuit specifications and must produce synthesisable Verilog HDL — evaluated across 50 tasks spanning combinational logic, sequential circuits, FSMs, and complex mixed designs. The framework goes beyond syntax checking: a semantic repair loop feeds simulation failures back to the model with structured error context, enabling iterative self-correction.

Conducted across 12 benchmark iterations — from a 5-task initial test to the full 50-task scope — evolving the pipeline from basic few-shot prompting through constrained prompts, post-processing, and finally semantic-aware iterative refinement with waveform analysis and formal equivalence checking.

Evaluation pipeline

Generate → simulate → repair → repeat

TASK SPEC 50 tasks 4 categories LLM Llama3 8B StarCoder2 7B TinyLlama 1.1B VERILOG HDL output post-processed SIMULATOR Verilator Icarus Verilog 5-METRIC SCORE SV · FC · GT waveform · formal equiv. check PASS ✓ results logged FAIL SEMANTIC REPAIR LOOP error context extracted → structured feedback → re-prompt LLM → re-simulate up to 6 iterations
How the repair loop works

Structured error context, not raw logs

When the simulator rejects generated Verilog, the naive approach is to dump the raw error log back to the model. This project takes a different approach: a feedback generator extracts structured error context — identifying whether the failure was a syntax issue, a port spec mismatch, an FSM state encoding problem, or a functional simulation failure — and formats it as targeted repair instructions rather than raw compiler output.

The pipeline evolved across 12 benchmark iterations. Phase 2 introduced constrained prompts with exact module and port name specifications, which alone produced a FSM breakthrough: StarCoder2 went from 0% to 66.7% syntax validity on FSM tasks. Phase 4 added AST-based repair and waveform diffing. Phase 5 introduced category-aware prompt templates for FSM and mixed designs.

The dataset is intentionally capped at 50 curated tasks — 23 combinational, 14 sequential, 8 FSM, 5 mixed — prioritising depth and analysis quality over breadth. Docker container enables full reproducibility on any platform.

# feedback_generator.py — structured error extraction

def generate_feedback(task, verilog, sim_result):
    errors = sim_result['errors']
    category = task['category']

    # Identify failure mode
    if 'syntax error' in errors:
        return build_syntax_feedback(errors, task)

    elif 'port mismatch' in errors:
        expected = task['inputs'] + task['outputs']
        return build_port_feedback(errors, expected)

    elif category == 'fsm':
        # Category-aware: FSM state encoding issues
        return build_fsm_feedback(
            errors, sim_result['waveform_diff']
        )

    return build_functional_feedback(
        errors, sim_result['expected_vs_actual']
    )
Results

Benchmark 12 — final 50-task results

Syntax validity and functional correctness across 3 models after Phase 5 (strict mode, waveform + formal enabled, up to 6 repair iterations per task).

Llama3 Syntax
71.3%
StarCoder2 Syntax
65%
TinyLlama Syntax
65%
Best Functional FC
51.8%
FSM (StarCoder2)
66.7%
Gen time reduction
−46%

Evolution across 12 benchmarks — functional correctness improvement at key phases.

PhaseMethodologyPass Rate
Phase 1Few-shot prompting~0%
Phase 2Constrained prompts + post-processing~25%
Phase 4Semantic repair + waveform analysis~42%
Phase 5Category-aware templates + micro-repair51.8%

Functional correctness on FSM tasks remains the hardest category — state encoding ambiguity is a fundamental LLM limitation for hardware description, not a prompt engineering problem. The paper discusses this as an open problem.

Full pipeline · Docker · 12 benchmark reports

Run the framework on any platform