评测科学视频生成的推理与知识能力,发现视觉真实不等于科学正确。
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

- 构建60个学科的1253个专家标注视频数据集,强调科学推理与知识融合。
- 模型在科学与因果正确性上表现差异大,开源模型显著落后于闭源模型。
- 提出可复现的评估框架,非专家与AI评判者也能接近专家意见。
我们提出Sci-VBench,一个涵盖自然科学、医疗健康、人文社科和工程四大领域共60个主题的综合性基准,包含1253个专家标注的视频生成任务。每个任务要求模型生成具有时序丰富性、依赖科学推理与知识约束的视频,超越表面视觉合理性。我们建立基于评分标准的评估协议,分析显示非专家人类评估者与多模态大模型作为评判者均能与专家判断保持较高一致性,支持大规模可复现评估。对16个前沿开源与闭源模型的基准测试表明,尽管自动感知质量评分在各系统间高度集中,但提示词对齐、科学与因果正确性表现差异显著,且存在明显的闭源-开源差距。结果表明,视觉真实性的进展尚未转化为对科学与因果动态的可靠建模。
原文摘要 · Abstract (English)
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。