Event date · · NetlistBench

NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation

FACT STATEMENT

NetlistBench is a structure-verified benchmark for SPICE netlist recognition and manipulation, containing 2,342 cases across 24 task families. It evaluates six non-thinking LLMs using a deterministic structure-aware oracle. Simple local edits achieve 96%-100% accuracy, device addition 41%-83%, and equivalence judgment 49%-90%. Enabling reasoning improves weaker models but does not eliminate structure-preservation failures, with performance degrading as edit horizon increases.

What happened

Researchers introduced NetlistBench, a benchmark with 2,342 cases across 24 task families to evaluate LLM reliability in SPICE netlist recognition and manipulation. Six non-thinking LLMs were tested using a deterministic structure-aware oracle. Accuracy varied by task complexity: simple local edits reached 96%-100%, device addition 41%-83%, and equivalence judgment 49%-90%. Enabling reasoning improved weaker models but did not eliminate structure-preservation failures, and performance degraded sharply with longer edit horizons.

Technical significance

The benchmark separates netlist-level structural manipulation from high-level design reasoning, using a deterministic structure-aware oracle to verify outputs. Performance correlates with operation-level structural complexity, and reasoning-enabled models still exhibit structure-preservation failures, indicating current LLMs lack robust internal representations of circuit topology and parameter constraints.

Industry impact

LLM adoption in circuit design workflows may be premature for tasks requiring precise netlist manipulation, such as device addition or equivalence checking. The observed accuracy gaps suggest a need for hybrid approaches combining LLMs with rule-based or formal verification tools before deployment in EDA pipelines.

Decision value

NetlistBench provides a standardized evaluation framework for companies developing AI-assisted circuit design tools, helping identify reliability gaps and guide investment in model improvement or hybrid systems.

What to watch

Expect follow-up work on improving LLM structural reasoning for netlists, possibly through fine-tuning on circuit-specific data or integration with symbolic solvers. Benchmark results may drive development of specialized models or guardrails for EDA applications.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.