arXiv:2608.16829cs.LGcs.AI2026-08被引 2

提出新基准测试视频模型对物理随机性的校准能力

CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?

论文配图:CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?
图 1 · 摘自论文原文
  • 用可解释的离散空间直接测量生成结果与真实分布距离
  • 发现多数模型概率集中于少数结果,显著偏离真实分布
  • 提供可复现的测试协议和量化指标,适合评估视频生成模型

视频世界模型通过生成采样近似物理结果的随机分布,但现有基准仅评分单个生成结果或粗略比较全数据集分布,未能检验特定现象的细粒度随机性。我们提出CaliBench,将结果评分置于可物理解释的离散空间(如格子编号、骰子点数、花色、颜色),而非学习特征空间(如FID),直接测量与已知参考分布的距离。我们构建了具有闭式参考分布的结果空间(二项式高尔顿板、伯努利分叉、均匀骰子/卡牌/彩票、偏斜的欧洲轮盘颜色),实现精确校准测试。我们将性能分解为两个正交维度:可评分性(生成结果中可评分的比例)与校准性(样本上与参考分布的总变差距离)。使用卡方检验评估显著性;因校准是零假设,仅能证明偏差,且在每单元N=32时仅能检测大偏差。我们在九个场景和六种图像到视频模型(WAN-2.7, SeeDance-2.0, HappyHorse-1.0, Veo 3.1, Runway Gen-4.5, Cosmos3-Super)上各生成32次。结果显示模型普遍将概率集中在少数结果,未再现参考分布。大多数场景-模型组合显著失准,极端情况如Veo 3.1在骰子场景中坍缩至单一结果。在轮盘场景中,部分生成结果使球位置模糊,导致多个模型可评分性低。性能随场景变化,无模型在所有场景中占优。我们发布测试协议和新指标(均归一化总变差,mnTV),供后续模型对比。

原文摘要 · Abstract (English)

Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score individual generations or compare distributions coarsely over a whole dataset, leaving the fine-grained aleatoric uncertainty of specific phenomena untested. We introduce CaliBench, which scores outcomes in a physically interpretable discrete space - a bin index, a die face, a suit, a colour - rather than a learned feature space such as in FID, so the distance from a known reference distribution is measured directly. We curate outcome spaces whose reference is known in closed form (binomial Galton boards, Bernoulli forks, uniform dice/cards/lottery, a skewed European-roulette colour), enabling an exact calibration test. We decompose performance into two orthogonal axes that a single accuracy metric conflates: scorability, the fraction of generations yielding a scoreable outcome, and calibration, the total variation distance from the reference on that sample. A chi-squared test assesses significance; as calibration is its null hypothesis it can evidence only miscalibration, and at N=32 per cell detects only large deviations. We apply it to nine scenes and six image-to-video models (WAN-2.7, SeeDance-2.0, HappyHorse-1.0, Veo 3.1, Runway Gen-4.5, Cosmos3-Super), 32 generations each. Models consistently concentrate probability mass on a few outcomes rather than reproducing the reference. Most scene-model combinations are significantly miscalibrated, in the extreme collapsing to one outcome, as Veo 3.1 does on dice. On roulette, generations often leave the ball ambiguously placed, giving several models low scorability. Performance varies by scene: no model dominates all nine. We release the protocol and a metric (mean normalised total variation, mnTV) for comparing new models against our results.

视频生成随机建模模型评估校准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。