arXiv:2503.21821cs.AI2025-03ACL被引 40

评测大模型解大学物理题能力,发现顶尖模型准确率仅59.9%

PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving

  • 构建1297道专家标注的大学物理题基准测试集
  • 最先进模型o3-mini在复杂物理问题上准确率仅59.9%
  • 提出RAG增强与提示工程策略,助力未来模型优化

我们提出PHYSICS,一个面向大学物理问题求解的综合性基准测试。该数据集包含1297道由专家标注的题目,覆盖经典力学、量子力学、热力学与统计物理、电磁学、原子物理和光学六大核心领域,每道题均需高级物理知识与数学推理能力。我们开发了可靠的自动化评估系统,用于精确验证模型表现。对主流基础模型的评估显示其存在显著局限:即使最先进的o3-mini模型,准确率也仅为59.9%,暴露出解决高水平科学问题的巨大挑战。通过全面的错误分析、多样提示策略探索以及基于检索增强生成(RAG)的知识增强方法,我们识别出关键改进方向,为后续研究奠定基础。

原文摘要 · Abstract (English)

We introduce PHYSICS, a comprehensive benchmark for university-level physics problem solving. It contains 1297 expert-annotated problems covering six core areas: classical mechanics, quantum mechanics, thermodynamics and statistical mechanics, electromagnetism, atomic physics, and optics. Each problem requires advanced physics knowledge and mathematical reasoning. We develop a robust automated evaluation system for precise and reliable validation. Our evaluation of leading foundation models reveals substantial limitations. Even the most advanced model, o3-mini, achieves only 59.9% accuracy, highlighting significant challenges in solving high-level scientific problems. Through comprehensive error analysis, exploration of diverse prompting strategies, and Retrieval-Augmented Generation (RAG)-based knowledge augmentation, we identify key areas for improvement, laying the foundation for future advancements.

物理推理大模型评测RAG科学计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。