评测大模型解大学物理题能力,发现顶尖模型准确率仅59.9%
PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving
- 构建1297道专家标注的大学物理题基准测试集
- 最先进模型o3-mini在复杂物理问题上准确率仅59.9%
- 提出RAG增强与提示工程策略,助力未来模型优化
我们提出PHYSICS,一个面向大学物理问题求解的综合性基准测试。该数据集包含1297道由专家标注的题目,覆盖经典力学、量子力学、热力学与统计物理、电磁学、原子物理和光学六大核心领域,每道题均需高级物理知识与数学推理能力。我们开发了可靠的自动化评估系统,用于精确验证模型表现。对主流基础模型的评估显示其存在显著局限:即使最先进的o3-mini模型,准确率也仅为59.9%,暴露出解决高水平科学问题的巨大挑战。通过全面的错误分析、多样提示策略探索以及基于检索增强生成(RAG)的知识增强方法,我们识别出关键改进方向,为后续研究奠定基础。
原文摘要 · Abstract (English)
We introduce PHYSICS, a comprehensive benchmark for university-level physics problem solving. It contains 1297 expert-annotated problems covering six core areas: classical mechanics, quantum mechanics, thermodynamics and statistical mechanics, electromagnetism, atomic physics, and optics. Each problem requires advanced physics knowledge and mathematical reasoning. We develop a robust automated evaluation system for precise and reliable validation. Our evaluation of leading foundation models reveals substantial limitations. Even the most advanced model, o3-mini, achieves only 59.9% accuracy, highlighting significant challenges in solving high-level scientific problems. Through comprehensive error analysis, exploration of diverse prompting strategies, and Retrieval-Augmented Generation (RAG)-based knowledge augmentation, we identify key areas for improvement, laying the foundation for future advancements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。