arXiv:2512.02942cs.CVcs.AI2025-12被引 4

首个评估视频生成模型科学推理能力的基准,检验其对物理化学现象的理解。

Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench

  • 设计复合科学场景,要求模型跨概念推理生成正确现象。
  • 在200个提示上测试7个主流模型,发现与人工评估高度一致。
  • 适合关注视频生成逻辑合理性与科学可信度的研究者使用。

视频生成的未来在于具备零样本推理能力的模型,理解真实世界科学规律对准确建模物理结果至关重要。然而现有视频基准多基于常识,难以反映模型的科学推理能力。我们提出VideoScience-Bench,一个用于评估本科级别科学理解能力的基准。每个提示包含需融合多个科学概念才能正确推断的复合场景,涵盖物理学与化学共14个主题、103个概念,共200个精心设计的提示。我们在七种先进视频生成模型(T2V和I2V)中进行专家标注评估,从五维指标:提示一致性、现象吻合度、正确动态性、不可变性、时空连续性展开分析。采用视觉语言模型作为评判器(VLM-as-a-Judge),结果显示其评估结果与人工评估高度相关。据我们所知,VideoScience-Bench是首个将视频生成模型同时视为生成者与推理者的评测基准,要求生成内容符合预期物理与化学现象。数据与代码已开源:github.com/hao-ai-lab/VideoScience。

原文摘要 · Abstract (English)

The next frontier for video generation lies in developing models capable of zero-shot reasoning, where understanding real-world scientific laws is crucial for accurate physical outcome modeling under diverse conditions. However, existing video benchmarks are physical commonsense-based, offering limited insight into video models' scientific reasoning capability. We introduce VideoScience-Bench, a benchmark designed to evaluate undergraduate-level scientific understanding in video models. Each prompt encodes a composite scientific scenario that requires understanding and reasoning across multiple scientific concepts to generate the correct phenomenon. The benchmark comprises 200 carefully curated prompts spanning 14 topics and 103 concepts in physics and chemistry. We conduct expert-annotated evaluations across seven state-of-the-art video models in T2V and I2V settings along five dimensions: Prompt Consistency, Phenomenon Congruency, Correct Dynamism, Immutability, and Spatio-Temporal Continuity. Using a VLM-as-a-Judge to assess video generations, we observe strong correlation with human assessments. To the best of our knowledge, VideoScience-Bench is the first benchmark to evaluate video models not only as generators but also as reasoners, requiring their generations to demonstrate scientific understanding consistent with expected physical and chemical phenomena. Our data and evaluation code are available at: \href{https://github.com/hao-ai-lab/VideoScience}{github.com/hao-ai-lab/VideoScience}.

视频生成科学推理基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。