让AI看懂物理竞赛图,实现视觉与科学推理的深度融合。
P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads
- 结合课程强化学习与智能体自检,提升模型推理稳定性。
- 在13场物理奥赛中斩获12金,开源模型表现领先全球第二。
- 适用于高阶科学推理,适合追求物理智能的科研与教育场景。
从符号运算到科学级推理的转变是大语言模型的关键挑战,而物理领域正是检验抽象逻辑与现实一致性的重要试金石。物理问题要求模型遵循宇宙基本规律,这需依赖多模态感知将抽象推理锚定于真实世界。在奥赛级别,图表常为关键约束来源,如边界条件和空间对称性,文本中往往缺失。为此,我们提出P1-VL,一系列专为高级科学推理设计的开源视觉-语言模型。方法融合课程强化学习(逐步提升难度以稳定后训练)与智能体增强(推理时迭代自验证)。在涵盖2024-2025年13场考试的HiPhO基准测试中,旗舰模型P1-VL-235B-A22B成为首个获12枚金牌的开源视觉-语言模型,并在开源模型中达到最先进水平。其智能体增强系统在全球排名第二,仅次于Gemini-3-Pro。此外,P1-VL在其他STEM基准上也显著优于基础模型,展现出强大泛化能力。通过开源,我们为通用物理智能提供基础,推动机器科学发现中视觉感知与物理定律的对齐。
原文摘要 · Abstract (English)
The transition from symbolic manipulation to science-grade reasoning represents a pivotal frontier for Large Language Models (LLMs), with physics serving as the critical test anchor for binding abstract logic to physical reality. Physics demands that a model maintain physical consistency with the laws governing the universe, a task that fundamentally requires multimodal perception to ground abstract logic in reality. At the Olympiad level, diagrams are often constitutive rather than illustrative, containing essential constraints, such as boundary conditions and spatial symmetries, that are absent from the text. To bridge this visual-logical gap, we introduce P1-VL, a family of open-source vision-language models engineered for advanced scientific reasoning. Our method harmonizes Curriculum Reinforcement Learning, which employs progressive difficulty expansion to stabilize post-training, with Agentic Augmentation, enabling iterative self-verification at inference. Evaluated on HiPhO, a rigorous benchmark of 13 exams from 2024-2025, our flagship P1-VL-235B-A22B becomes the first open-source Vision-Language Model (VLM) to secure 12 gold medals and achieves the state-of-the-art performance in the open-source models. Our agent-augmented system achieves the No.2 overall rank globally, trailing only Gemini-3-Pro. Beyond physics, P1-VL demonstrates remarkable scientific reasoning capacity and generalizability, establishing significant leads over base models in STEM benchmarks. By open-sourcing P1-VL, we provide a foundational step toward general-purpose physical intelligence to better align visual perceptions with abstract physical laws for machine scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。