arXiv:2602.15037cs.SEcs.AI2026-02被引 10

测试大模型在电路分析中遵守规则与真实理解能力的差距

CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis

  • 设计对照实验对齐规则遵循与物理理解能力
  • 强模型解题准确但常违背符号约定,弱模型更听话却易出错
  • 适合研究AI对齐、工程教育及高可靠性系统评估

随着大语言模型在工程领域接近专家水平,用户指定约束下的可靠推理变得至关重要。以电路分析为例,即使数值正确,若违反回路方向或极性等方法论规范,仍可能在安全关键系统中引发错误传播。然而,当前前沿模型是基于第一性原理推理,还是依赖与指令冲突的训练先验尚不明确。本文提出CircuChain诊断基准,旨在分离电路分析中的指令合规性与物理推理能力。该基准包含五种典型电路拓扑的平衡对照组(控制/陷阱问题对),并系统改变符号约定、电流方向和极性定义。通过结合符号求解器、SPICE仿真及基于LLM的错误分类体系,实现故障来源的细粒度归因:规则错误、物理错误、算术失误或幻觉。在每模型100个任务中,均观察到‘合规-能力分化’现象:最强模型具备近乎完美的物理推理能力,但在故意反转自然符号模式的陷阱条件下表现出高违规率;而较弱模型虽物理准确性较低,却更严格遵守显式指令。结果表明,模型能力提升并不保证约束对齐,凸显需构建数学严格领域的指令遵循评估框架。CircuChain为此提供了一种可行方案,并为工程教育与AI对齐研究提供可操作洞见。

原文摘要 · Abstract (English)

As large language models (LLMs) advance toward expert-level performance in engineering domains, reliable reasoning under user-specified constraints becomes critical. In circuit analysis, for example, a numerically correct solution is insufficient if it violates established methodological conventions such as mesh directionality or polarity assignments, errors that can propagate in safety-critical systems. Yet it remains unclear whether frontier models truly apply first-principles reasoning or rely on entrenched training priors that conflict with explicit instructions. We introduce CircuChain, a diagnostic benchmark designed to disentangle instruction compliance from physical reasoning competence in electrical circuit analysis. CircuChain consists of counterbalanced Control/Trap problem pairs across five canonical circuit topologies, augmented with systematic variations in sign conventions, current orientations, and polarity definitions. A multi-stage verification pipeline, combining symbolic solvers, SPICE simulation, and an LLM-based error taxonomy, enables fine-grained attribution of failures to convention errors, physics errors, arithmetic mistakes, or hallucinations. Across 100 tasks per model, we observe a consistent Compliance-Competence Divergence. The strongest model evaluated exhibits near-perfect physical reasoning but a high rate of convention violations when Trap conditions deliberately invert natural sign patterns. Conversely, weaker models display lower physical fidelity yet superior adherence to explicit instructions. These results suggest that increased model capability does not guarantee improved constraint alignment and highlight the need for new evaluation frameworks that stress instruction-following under mathematically rigid domains. CircuChain provides one such framework and offers actionable insights for both engineering education and AI alignment research.

大模型评估电路分析指令对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。