arXiv:2607.00276cs.LGcs.AI2026-07

测试大模型在陌生物理世界中的真实推理能力,发现其常凭记忆而非逻辑判断。

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds

论文配图:Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds
图 1 · 摘自论文原文
  • 设计四阶段诊断流程,通过诱导、建模、预测与自检评估模型真实推理能力
  • 三类陌生物理框架下,顶级模型平均仅6/15通过测试,且多因定量计算失误
  • 揭示模型自审能力弱、跨框架判断不可靠,适合关注AI推理可信度的研究者

当前大语言模型(LLM)的物理能力评测主要依赖答对率,难以区分真实推理与模式记忆,也无法定位推理失败点。本文提出可审计的四阶段诊断方法,评估模型在陌生物理框架中是否具备归纳、建模、预测与复盘能力。该方法结合预注册锁定、阶段间独立会话、双模型评判与人工审核路径,应用于三种平行物理世界:单方程反事实世界(F=mv)、历史框架(亚里士多德力学)与四域反事实世界(衰变世界)。在Claude Opus 4.7、GPT-5.5和Gemini 3.1 Pro中,三类世界的综合通过率分别为6/15、6/15、0/15(其中F=mv与亚里士多德力学需内容与结构双重正确,衰变世界仅内容轴有效)。最显著现象为定性与定量不对称:模型极少预测错误变化方向,但频繁因沿用标准物理关系而计算错误比例。此外,模型评判可靠性不跨框架迁移,且第四阶段自检普遍失效——至少三分之二存在错误的实验中,模型自我报告未出错。所有提示、响应、判别与审计记录均已开源。

原文摘要 · Abstract (English)

Current large-language-model (LLM) physics benchmarks are usually scored by answer accuracy, which cannot distinguish genuine reasoning from recall of familiar problem patterns and reveals little about where a model's reasoning breaks down. We introduce an auditable four-stage diagnostic that evaluates whether an LLM can reason inside an unfamiliar physics framework through induction, formulation, prediction, and review. The diagnostic combines locked pre-registrations, fresh sessions between stages, dual-LLM judging, and a human-audit pathway, and we apply it to three parallel physics worlds: a single-equation counterfactual world ($F=mv$), a historical framework (Aristotelian mechanics), and a four-domain counterfactual world (Decay World). Across Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro, the three worlds yield composite PASS rates are 6/15, 6/15, and 0/15 respectively (content $\land$ structural for $F=mv$ and Aristotelian, content axis only for Decay World where the structural axis is out of scope). The most pointed empirical pattern is a qualitative-versus-quantitative asymmetry: in Decay World, models almost never predict the wrong direction of change, but frequently compute the wrong ratio by slipping back to standard-physics relations. The protocol also surfaces two methodology findings: LLM-judge reliability does not transfer across frameworks, and Stage 4 self-review is weak in every framework, with the model's own review wrongly reporting no earlier error in at least two-thirds of the trials that actually contained one. We release the full prompts, responses, verdicts, and audit records.

大模型推理物理认知可审计评估AI可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。