新基准测试发现视频模型在复杂推理上严重不足。
How Far Are Video Models from True Multimodal Reasoning?

- 设计新评测框架CLVG-Bench,用上下文学习测试零样本推理能力。
- SOTA模型在逻辑推理任务成功率低于25%,交互生成几乎失败。
- 适合关注视频生成与多模态推理的开发者和研究者参考。
尽管通用视频模型取得显著进展,但一个关键问题仍未解决:这些模型距离真正多模态推理还有多远?现有基准因任务设计简单、评估指标分散,无法严谨回答此问题。为此,我们提出CLVG-Bench,一个通过上下文学习视频生成来探测视频模型零样本推理能力的评测框架。该框架包含超过1000条高质量人工标注的元数据,覆盖6大类47个子类,涵盖物理模拟、逻辑推理和交互情境等复杂场景。为实现可扩展且可靠的评估,我们进一步提出自适应视频评估器AVE,仅用少量标注即可对齐人类专家判断,提供可解释的文本反馈。大量实验揭示:尽管顶尖模型如Seedance 2.0在部分理解与推理任务中表现良好,但在基于逻辑的生成任务(成功率<25%)和交互生成任务(成功率~0%)上严重不足,暴露出多模态推理与物理一致性是主要瓶颈。该方法系统量化了局限性,提供可操作反馈与迈向鲁棒通用视频模型的路线图。CLVG-Bench与代码已公开。
原文摘要 · Abstract (English)
Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true multimodal reasoning? Existing benchmarks fail to address this question rigorously, as they remain constrained by straightforward task designs and fragmented evaluation metrics that neglect complex multimodal reasoning. To bridge this gap, we introduce CLVG-Bench, an evaluation framework designed to probe video models' zero-shot reasoning capabilities via Context Learning in Video Generation. CLVG-Bench comprises more than 1,000 high-quality, manually annotated metadata across 6 categories and 47 subcategories, covering complex scenarios including physical simulation, logical reasoning, and interactive contexts. To enable rigorous and scalable assessment, we further propose an Adaptive Video Evaluator (AVE) that aligns with human expert perception using minimal annotations, delivering interpretable textual feedback across diverse video context tasks. Extensive experiments reveal a striking answer to our central question: while state-of-the-art (SOTA) video models, such as Seedance 2.0, demonstrate competence on certain understanding and reasoning subtasks, they fall substantially short with logically grounded and interactive generation tasks (achieving success rates <25% and ~0%, respectively), exposing multimodal reasoning and physical grounding as critical bottlenecks. By systematically quantifying these limitations, the proposed method provides actionable feedbacks and a clear roadmap toward truly robust, general-purpose video models. CLVG-Bench and code are released here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。