用强化学习训练出能拿物理奥赛金牌的开源大模型。
P1: Mastering Physics Olympiads with Reinforcement Learning
- 全靠强化学习训练,专攻物理奥赛级推理
- P1-235B-A22B在2025年国际物理奥赛中获金牌,13项赛事12金
- 模型通用性强,数学编程也表现优异
大型语言模型(LLMs)的进展已从解谜转向科学级推理——这种推理需经得起自然规律检验,而非仅符合评分标准。物理是此类能力的最严苛测试,因其将符号与现实紧密绑定,是现代技术的基石。本文提出P1系列开源物理推理模型,全部通过强化学习(RL)训练。其中,P1-235B-A22B是首个在2025年国际物理奥林匹克竞赛(IPhO 2025)中取得金牌成绩的开源模型,并在2024/2025年度13项国际/地区性物理竞赛中斩获12枚金牌。P1-30B-A3B在该竞赛中亦获得银牌,超越几乎所有其他开源模型。进一步结合代理框架PhysicsMinions后,P1-235B-A22B+PhysicsMinions在IPhO 2025中总分排名第一,平均得分最高。此外,P1系列在数学与编程等推理任务上也表现卓越,展现出极强的泛化能力。
原文摘要 · Abstract (English)
Recent progress in large language models (LLMs) has moved the frontier from puzzle-solving to science-grade reasoning-the kind needed to tackle problems whose answers must stand against nature, not merely fit a rubric. Physics is the sharpest test of this shift, which binds symbols to reality in a fundamental way, serving as the cornerstone of most modern technologies. In this work, we manage to advance physics research by developing large language models with exceptional physics reasoning capabilities, especially excel at solving Olympiad-level physics problems. We introduce P1, a family of open-source physics reasoning models trained entirely through reinforcement learning (RL). Among them, P1-235B-A22B is the first open-source model with Gold-medal performance at the latest International Physics Olympiad (IPhO 2025), and wins 12 gold medals out of 13 international/regional physics competitions in 2024/2025. P1-30B-A3B also surpasses almost all other open-source models on IPhO 2025, getting a silver medal. Further equipped with an agentic framework PhysicsMinions, P1-235B-A22B+PhysicsMinions achieves overall No.1 on IPhO 2025, and obtains the highest average score over the 13 physics competitions. Besides physics, P1 models also present great performance on other reasoning tasks like math and coding, showing the great generalibility of P1 series.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。