首个多模态物理推理基准,评估模型三方面能力。
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
- 从变量识别、过程推导到解题全程评估多模态模型
- 涵盖图文输入的物理问题理解与推理全流程
- 适合研究多模态推理与智能教育系统的开发者
多模态大语言模型在多样化推理任务中表现出色,但在复杂物理推理领域的应用仍不充分。物理推理面临独特挑战:需基于物理情境并整合多模态信息。现有基准多局限于纯文本输入或仅关注求解结果,忽视变量识别与过程构建等关键中间步骤。为此,我们提出PhysicsArena,首个全面覆盖三个核心维度的多模态物理推理基准:变量识别、物理过程建模与解题推导。该基准旨在为评估和提升多模态大模型在物理推理方面的能力提供综合性平台。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in diverse reasoning tasks, yet their application to complex physics reasoning remains underexplored. Physics reasoning presents unique challenges, requiring grounding in physical conditions and the interpretation of multimodal information. Current physics benchmarks are limited, often focusing on text-only inputs or solely on problem-solving, thereby overlooking the critical intermediate steps of variable identification and process formulation. To address these limitations, we introduce PhysicsArena, the first multimodal physics reasoning benchmark designed to holistically evaluate MLLMs across three critical dimensions: variable identification, physical process formulation, and solution derivation. PhysicsArena aims to provide a comprehensive platform for assessing and advancing the multimodal physics reasoning abilities of MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。