arXiv:2605.30346cs.CV2026-05被引 1

用反向视频测试视频生成模型是否真懂因果关系

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

论文配图:YoCausal: How Far is Video Generation from World Model? A Causality Perspective
图 1 · 摘自论文原文
  • 通过反转真实视频构建反事实样本,评估模型对时间方向的感知
  • 13个主流视频生成模型在因果认知上表现远低于人类水平
  • 提出双层级评测框架,区分时间模式与真实因果推理

随着视频扩散模型(VDMs)向世界模型演进,一个核心问题浮现:它们是否真正理解因果关系,还是仅依赖统计时间模式?现有基准多基于合成数据,因模拟到现实的差距而限制了真实泛化能力。我们提出YoCausal,一种受认知科学中预期违背(VoE)范式启发的两级评估基准。通过零成本反转真实视频生成自然反事实样本,建立可任意扩展的评估协议。一级引入反向惊喜指数(RSI),以去噪损失量化时间箭头感知;二级引入因果认知指数(CCI),利用视觉语言模型将数据集划分为因果与非因果子集,分离真实因果推理与时间偏见。对13个顶尖VDM的评估显示,感知时间方向并不等同于理解因果,与人类水平的因果认知仍存在显著差距。

原文摘要 · Abstract (English)

As video diffusion models (VDMs) advance toward world models, a key question arises: do they truly understand causality, or merely overfit to statistical temporal patterns? Existing benchmarks mostly rely on synthetic data, limiting real-world generalization due to the sim-to-real gap. We present YoCausal, a two-level benchmark inspired by the Violation of Expectation (VoE) paradigm from cognitive science. By temporally reversing real-world videos at zero cost as natural counterfactual samples, YoCausal establishes an arbitrarily extensible evaluation protocol. Level 1 introduces the Reverse Surprise Index (RSI), quantifying arrow-of-time perception via denoising loss. Level 2 introduces the Causality Cognition Index (CCI), which leverages a VLM to stratify datasets into causal and non-causal subsets, disentangling genuine causal reasoning from temporal bias. Evaluation of 13 state-of-the-art VDMs reveals that perceiving the arrow of time does not imply understanding causality, and a significant gap persists relative to human-level causal cognition.

视频生成因果推理世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。