构建中文物理多模态基准,评估模型理解与生成能力
OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

- 基于中文教育语料构建覆盖中学到大学的多模态物理题库
- 包含1.5万题、2万张图,支持细粒度推理分析
- 首次系统评估模型生成物理图表的能力,适合科研与教学
多模态大语言模型在视觉与文本推理任务中表现优异,但在物理领域的发展受限于缺乏全面的评测基准。为此,我们提出OmniPhys,一个大规模多模态物理理解与推理基准,涵盖中国教育语料中的中学至大学水平题目。OmniPhys包含15,246道题目和19,850张图像,并配有详细标注,支持对推理过程与知识应用的细粒度分析。该基准不仅评估传统任务,还系统评估模型在物理领域的多模态输出能力,包括生成结构化物理图示的能力——这是真实物理问题求解的核心组成部分。大量实验揭示当前多模态大模型在复杂推理与视觉生成方面存在显著短板。为此,我们公开OmniPhys作为推动物理与科学领域多模态智能发展的基础资源。代码与数据已发布于https://github.com/ECNU-RAIL/OmniPhys-EMNLP2026。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of a comprehensive benchmark. To fill this gap, we introduce OmniPhys, a large-scale benchmark for multimodal physics understanding and reasoning, covering middle school through university-level problems from Chinese Educational Corpora. OmniPhys consists of 15,246 questions and 19,850 images, accompanied by detailed annotations that support fine-grained analysis of reasoning processes and knowledge usage. Beyond conventional evaluation, OmniPhys is a benchmark that systematically evaluates multimodal outputs in the physics domain, including models' ability to generate structured physics diagrams, which constitute a fundamental component of authentic physics problem solving. Extensive evaluations reveal critical gaps in the capabilities of current MLLMs, especially in complex reasoning and visual generation. To address this, we release OmniPhys to serve as a foundational resource for advancing multimodal intelligence in physics and scientific domains. Codes and data are available at https://github.com/ECNU-RAIL/OmniPhys-EMNLP2026.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。