arXiv:2604.14188cs.AIcs.CL2026-04被引 1

用五级评分体系评估大模型在高阶物理中的隐含推理能力

Grading the Unspoken: Evaluating Tacit Reasoning in Quantum Field Theory and String Theory with LLMs

  • 构建12道量子场论与弦论核心问题数据集,设计五级评分标准
  • 模型在显式推导中表现接近满分,但面对隐含推理时明显退化
  • 失败主因是概念框架选择不稳定,无法重构省略的推理步骤

大语言模型在数学与物理多个领域表现出色。一个自然问题是:它们能否支持量子场论和弦论这类高度抽象理论领域的研究?评估这一可能性面临挑战:这些领域的正确性具有层次性、隐含性,且本质上非二元。传统答案匹配指标无法捕捉中间概念步骤是否合理重建,或隐式结构约束是否被遵守。我们构建了一个由12个问题组成的专家精选数据集,覆盖量子场论与弦论核心内容,并提出五级评分标准,区分陈述正确性、关键概念意识、推理链存在性、隐含步骤重构和内容拓展性。评估多个主流大模型发现,在稳定概念框架内的显式推导中表现接近上限,但在需要重构省略推理步骤或在全局一致性约束下重组表示时系统性退化。这些失败不仅源于缺失中间步骤,更源于表示选择的不稳定性:模型常无法识别解决隐式矛盾所需的正确概念框架。我们认为,高度抽象的理论物理为当前评估范式的认识论局限提供了一个极为敏感的观测窗口。

原文摘要 · Abstract (English)

Large language models have demonstrated impressive performance across many domains of mathematics and physics. One natural question is whether such models can support research in highly abstract theoretical fields such as quantum field theory and string theory. Evaluating this possibility faces an immediate challenge: correctness in these domains is layered, tacit, and fundamentally non-binary. Standard answer-matching metrics fail to capture whether intermediate conceptual steps are properly reconstructed or whether implicit structural constraints are respected. We construct a compact expert-curated dataset of twelve questions spanning core areas of quantum field theory and string theory, and introduce a five-level grading rubric separating statement correctness, key concept awareness, reasoning chain presence, tacit step reconstruction, and enrichment. Evaluating multiple contemporary LLMs, we observe near-ceiling performance on explicit derivations within stable conceptual frames, but systematic degradation when tasks require reconstruction of omitted reasoning steps or reorganization of representations under global consistency constraints. These failures are driven not only by missing intermediate steps, but by an instability in representation selection: models often fail to identify the correct conceptual framing required to resolve implicit tensions. We argue that highly abstract theoretical physics provides a uniquely sensitive lens on the epistemic limits of current evaluation paradigms.

大模型评估理论物理隐含推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。