评测大模型物理推理能力,设计了500道原创难题
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
- 构建500道原创物理题,避免数据污染
- 最佳模型仅36.9%准确率,远低于人类61.9%
- 引入表达式编辑距离提升评估效率204%
当前大语言模型推理能力评估基准存在任务简化、数据污染和评估项缺陷等问题。为此,我们提出PHYBench,一个包含500道原创物理题的基准,难度涵盖高中到物理奥赛级别。通过系统化筛选流程,确保内容原创并剔除错误题目。评估显示,该基准激活更多推理令牌,对不同模型的区分度优于AIME 2024、OlympiadBench和GPQA。即使最优模型Gemini 2.5 Pro,准确率也仅为36.9%,远低于人类专家的61.9%。为提升数学表达评估精度,我们引入表达式编辑距离(EED)评分,相较二值评分可提升样本效率204%。PHYBench能有效激发多步、多条件推理,可用于考察模型推理鲁棒性、偏好与缺陷。数据集及评测结果已公开于https://www.phybench.cn/。
原文摘要 · Abstract (English)
Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies necessitate more rigorous assessment methods. To address these limitations, we introduce PHYBench, a benchmark of 500 original physics problems ranging from high school to Physics Olympiad difficulty. PHYBench addresses data contamination through original content and employs a systematic curation pipeline to eliminate flawed items. Evaluations show that PHYBench activates more tokens and provides stronger differentiation between reasoning models compared to other baselines like AIME 2024, OlympiadBench and GPQA. Even the best-performing model, Gemini 2.5 Pro, achieves only 36.9% accuracy compared to human experts' 61.9%. To further enhance evaluation precision, we introduce the Expression Edit Distance (EED) Score for mathematical expression assessment, which improves sample efficiency by 204% over binary scoring. Moreover, PHYBench effectively elicits multi-step and multi-condition reasoning, providing a platform for examining models' reasoning robustness, preferences, and deficiencies. The benchmark results and dataset are publicly available at https://www.phybench.cn/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。