arXiv:2510.11606cs.CV2025-10被引 2

首个评估科学实验视频理解的基准,揭示大模型在细节与推理上的短板。

ExpVid: A Benchmark for Experiment Video Understanding & Reasoning

  • 构建三层次任务体系:感知工具材料、理解操作顺序、推导科学结论。
  • 19个主流模型在细粒度识别和长期推理上表现差,开源模型差距明显。
  • 专为湿实验视频设计,适合研究多模态大模型在科研中的应用者。

多模态大语言模型(MLLMs)有望通过解析复杂实验流程加速科学发现。然而,现有基准未能反映真实实验室工作中的细粒度与长时程特性,尤其在湿实验场景中。为此,我们提出ExpVid,首个系统评估MLLMs在科学实验视频理解能力的基准。数据源自同行评审的视频论文,包含三层次任务体系:(1) 工具、材料与动作的细粒度感知;(2) 步骤顺序与完整性理解;(3) 将完整实验过程与已发表结论关联的科学推理。通过视觉中心标注流程,结合自动化生成与多领域专家验证,确保任务需视觉定位支撑。我们在ExpVid上评估了19个领先MLLMs,发现尽管在粗粒度识别上表现良好,但在消歧细部、追踪状态变化及连接实验步骤与科学结论方面存在显著困难。结果揭示了专有模型与开源模型在高阶推理上的显著性能差距。ExpVid不仅是诊断工具,也为发展可信赖的科研协作型大模型指明方向。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) hold promise for accelerating scientific discovery by interpreting complex experimental procedures. However, their true capabilities are poorly understood, as existing benchmarks neglect the fine-grained and long-horizon nature of authentic laboratory work, especially in wet-lab settings. To bridge this gap, we introduce ExpVid, the first benchmark designed to systematically evaluate MLLMs on scientific experiment videos. Curated from peer-reviewed video publications, ExpVid features a new three-level task hierarchy that mirrors the scientific process: (1) Fine-grained Perception of tools, materials, and actions; (2) Procedural Understanding of step order and completeness; and (3) Scientific Reasoning that connects the full experiment to its published conclusions. Our vision-centric annotation pipeline, combining automated generation with multi-disciplinary expert validation, ensures that tasks require visual grounding. We evaluate 19 leading MLLMs on ExpVid and find that while they excel at coarse-grained recognition, they struggle with disambiguating fine details, tracking state changes over time, and linking experimental procedures to scientific outcomes. Our results reveal a notable performance gap between proprietary and open-source models, particularly in high-order reasoning. ExpVid not only provides a diagnostic tool but also charts a roadmap for developing MLLMs capable of becoming trustworthy partners in scientific experimentation.

实验视频多模态科学推理大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。