arXiv:2608.09374cs.AIcs.CV2026-08

评测模型从电路图到符号推理的全程能力,发现现有模型在长链条推理中表现明显下降。

CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits

论文配图:CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits
图 1 · 摘自论文原文
  • 构建1000个真实课本题,涵盖图符识别、拓扑还原、方程建模等全流程
  • 多模型测试最高准确率84.8%,但长任务下性能显著下滑
  • 适合评估多模态大模型在工程推理中的物理一致性与逻辑持续性

电路分析不仅需要识别图像中的元件,还需对符号和标签进行语义映射,恢复隐含拓扑结构,选择物理模型,建立耦合方程,传递中间量,并保持单位、正负号、方向和相位约定。我们提出\benchmark,一个包含1000个真实教材问题的基准,用于评估这一完整的长时程视觉到符号推理过程。每个问题配有一张或多张电路图、自包含的问题、文本或语义指定的答案及参考解答。采用证据优先的构建流程对齐问题、图表与解答,基于推理导向的分类体系按电路类型和依赖深度组织题目。评估结合保守的类型化评分与身份盲目的多模型语义共识,保留所有题目作为分母。在三个商业聊天机器人和六个开源多模态大模型上,最高系统达到84.8%准确率。然而,性能在长时程问题上持续下降,定性分析揭示了拓扑到目标绑定、物理约定和后期输出传播方面的持久失败。\benchmark{}为测量多模态模型将技术视觉证据转化为持续且物理有效的符号推理提供了聚焦测试平台。代码已公开于GitHub - CircuitReason/CircuitReason1K。

原文摘要 · Abstract (English)

Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, select a physical model, formulate coupled equations, propagate intermediate quantities, and preserve units, signs, directions, and phase conventions. We introduce \benchmark, a benchmark of 1,000 authentic textbook problems for evaluating this complete long-horizon visual-to-symbolic reasoning process. Each problem pairs one or more circuit diagrams with a self-contained question, a typed or semantically specified answer, and a reference worked solution. An evidence-first construction pipeline aligns questions, figures, and solutions, while a reasoning-oriented taxonomy organizes problems by circuit type and dependency depth. Evaluation combines conservative typed scoring with identity-blinded multi-model semantic consensus, retaining every problem in the denominator. Across three commercial chatbot systems and six open-source multimodal large language models, the highest-scoring system reaches 84.8\% accuracy. However, performance consistently deteriorates on long-horizon problems, and qualitative analysis exposes persistent failures in topology-to-target binding, physical conventions, and late-stage output propagation. \benchmark{} provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbolic reasoning. Code are available at GitHub - CircuitReason/CircuitReason1K.

电路推理多模态模型长时程推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。