用高难度动态物理题测试大模型推理能力,发现其泛化短板。
ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems
- 构建静态+动态双模块物理题库,覆盖奥赛级难题
- 400道静态题+100道可自动变式的动态题,答案需精确格式
- 揭示大模型在动态变化场景中严重泛化不足
大型语言模型在数学和编程领域表现优异,但在物理推理方面仍缺乏深入探索与理解。物理问题要求精确计算、深层概念理解及物理建模能力,现有基准普遍因难度不足、采用选择题形式、评估环境静态而无法有效衡量模型的物理建模能力。本文提出ABench-Physics,一个全新基准,用于严格评估大模型的物理推理与泛化能力。该基准包含两部分:Phy_A(400道研究生或奥赛级别静态题)和Phy_B(100道配备自动变体引擎的动态题),所有题目均需精确数值回答,并设定严格的格式与容差要求。对多个前沿大模型的评估显示显著性能差距,凸显其在应对动态变化条件时的持续局限。ABench-Physics为推进大模型科学推理能力提供了挑战性且诊断性强的框架。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown impressive performance in domains such as mathematics and programming, yet their capabilities in physics remain underexplored and poorly understood. Physics poses unique challenges that demand not only precise computation but also deep conceptual understanding and physical modeling skills. Existing benchmarks often fall short due to limited difficulty, multiple-choice formats, and static evaluation settings that fail to capture physical modeling ability. In this paper, we introduce ABench-Physics, a novel benchmark designed to rigorously evaluate LLMs' physical reasoning and generalization capabilities. ABench-Physics consists of two components: Phy_A, a static set of 400 graduate- or Olympiad-level problems; and Phy_B, a dynamic subset of 100 problems equipped with an automatic variation engine to test model robustness across changing conditions. All questions require precise numerical answers, with strict formatting and tolerance constraints. Our evaluation of several state-of-the-art LLMs reveals substantial performance gaps, highlighting persistent limitations in physical reasoning, especially in generalization to dynamic variants. ABench-Physics provides a challenging and diagnostic framework for advancing scientific reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。