构建物理原则推理测试集,揭示大模型缺乏专家式简洁解题能力
PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models
- 设计基于物理原理的推理评测基准,引导专家用核心原理快速解题
- 多款主流大模型在该基准上表现不佳,无法复现专家解题路径
- 适合关注科学推理可解释性与模型本质理解的研究者使用
大型语言模型(LLMs)在科学问题求解方面进展迅速,尤其在物理领域表现出较强能力。然而,当前模型常无法模仿人类专家那种简洁、以原理为核心的推理方式,反而生成冗长且难以理解的解答。这一差距凸显了其在应用核心物理原理进行高效、可解释性求解方面的显著不足。为系统评估此局限性,我们提出 PhySense——一个基于物理原理的新型推理评测基准。该基准设计为专家可轻松通过核心原理快速求解,但对未采用‘先原理’思维的大模型而言却极具挑战。我们在多个前沿大模型及不同提示策略下进行评估,结果一致显示模型未能与专家式推理路径对齐,为开发具备高效、稳健、可解释性的原理驱动型科学推理人工智能系统提供了重要洞见。
原文摘要 · Abstract (English)
Large language models (LLMs) have rapidly advanced and are increasingly capable of tackling complex scientific problems, including those in physics. Despite this progress, current LLMs often fail to emulate the concise, principle-based reasoning characteristic of human experts, instead generating lengthy and opaque solutions. This discrepancy highlights a crucial gap in their ability to apply core physical principles for efficient and interpretable problem solving. To systematically investigate this limitation, we introduce PhySense, a novel principle-based physics reasoning benchmark designed to be easily solvable by experts using guiding principles, yet deceptively difficult for LLMs without principle-first reasoning. Our evaluation across multiple state-of-the-art LLMs and prompt types reveals a consistent failure to align with expert-like reasoning paths, providing insights for developing AI systems with efficient, robust and interpretable principle-based scientific reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。