测试大模型对物理图示的全局推理能力,发现其严重短板。
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning

- 构建2000+任务的费曼图基准,要求模型理解拓扑与代数关系。
- 模型局部识别准确率70-95%,但整体结构重建仅13-17%。
- 适合评估科学图示推理能力,尤其关注拓扑敏感任务。
当前多模态科学推理基准主要评估局部信息提取——模型识别符号和数值后进行文本推理。它们未检验模型是否能基于正式图示的全局结构特性(如拓扑、守恒约束、视觉模式与代数表达的一致映射)进行推理。我们提出FeynmanBench,一个包含2000多个任务的基准,聚焦标准模型中的电磁、弱、强相互作用的费曼图。每个任务结合图示与最小文本约定,要求模型恢复完整物理内容:顶点清单、传播子类型、拓扑连接性、动量路由及完整的散射振幅。自动化生成与验证流水线在标准化规则下生成图示、标注与参考答案。评估19个先进多模态大模型发现,模型在局部识别(顶点与传播子识别)上达70–95%,但在拓扑重构(CP3)上降至13–17%,全代数推导(CP5)几乎为零。FeynmanBench为形式化科学图示的多模态推理提供受控测试平台,并揭示现有架构在拓扑敏感科学推理中的根本局限。
原文摘要 · Abstract (English)
Current multimodal benchmarks for scientific reasoning primarily evaluate local information extraction -- models recognize symbols and values and then perform textual inference. They do not assess whether models can reason over the global structural properties of formal diagrams, such as topology, conservation constraints, and the consistent mapping between visual patterns and algebraic expressions. We introduce FeynmanBench, a benchmark of over 2,000 tasks centered on Feynman diagrams spanning the electromagnetic, weak, and strong interactions of the Standard Model. Each instance couples a diagram image with minimal textual conventions and requires models to recover the full physical content -- vertex inventory, propagator types, topological connectivity, momentum routing, and the complete scattering amplitude. An automated generation and verification pipeline produces the diagrams, annotations, and reference answers under standardized rules. Evaluating 19 state-of-the-art multimodal LLMs, we find a consistent failure pattern: models achieve 70--95\% on local recognition (vertex and propagator identification) but collapse to 13--17\% on topological reconstruction (CP3), and near zero on full algebraic derivation (CP5). FeynmanBench offers a controlled testbed for multimodal reasoning over formal scientific diagrams and highlights fundamental limitations of current architectures in topology-sensitive scientific reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。