视频生成模型难以模拟不可逆过程,实测发现其进展极弱,仅能维持静态。
Diagnosing Under-Development of Irreversible Processes in Video Generation

- 用进展度与静止率代替反向检测,更可靠地衡量不可逆性
- 真实视频进展率ρ=+0.40,生成视频几乎无进展,静止率高达92%-100%
- 揭示生成模型存在不可逆属性的“发育不足”,适合关注视频物理合理性研究者
许多物理属性具有不可逆性:冰融化后不会自复,纸烧焦后不会自还原。视频生成模型是否遵循这一规律?我们发现该问题难以测量,真正可靠的指标是“发展”而非“逆转”。局部逆转检测指标存在零假设缺陷:在纯噪声上,每片段违规率仍达0.50;方差归一化逆转残差也处于噪声上限。唯一通过零检验的是双阶段协议:进展(方向性属性相关性)与静止率。在此协议下,真实视频与生成视频清晰分离,且经人工验证。七种文本到视频模型中,真实参考视频进展ρ=+0.40,35%静止;所有生成器进展接近零,静止率92–100%;九名标注者评分显示真实视频(2.75)显著高于生成视频(0.99,0–4分制)。核心发现为“发育不足”:生成模型对不可逆属性的发展极其有限,而非逆转。此外,后处理读出引导易被操纵,而通过解耦属性潜在空间构建单调性约束则可避免此问题,在受控与半合成场景中得到验证。
原文摘要 · Abstract (English)
Many physical attributes are \emph{irreversible}: ice melts but does not re-freeze, paper chars but does not un-burn. Do video generators respect this? We show the question is hard to measure, and that what can be measured reliably is \emph{development} rather than reversal. Metrics of local reversal are null-degenerate: a per-clip violation rate scores $0.50$ on pure noise, and a variance-normalized reversal residual sits at its noise ceiling. What survives null-testing is a two-part protocol: progress (a directional attribute correlation) and a stasis rate. Under this protocol, generated video separates cleanly from real footage, and the gap is human-validated. Across seven text-to-video models, real reference footage advances ($ρ{=}{+}0.40$, $35\%$ static) while every generator shows near-zero progress and $92$--$100\%$ stasis; nine annotators rate real footage far above generated ($2.75$ vs.\ $0.99$ on a $0$--$4$ scale). The reliable finding is \emph{under-development}: generators barely advance irreversible attributes rather than reversing them. As a complementary mechanism, we show that post-hoc readout guidance is gameable, whereas enforcing monotonicity by construction in a disentangled attribute latent removes the gameable readout, validated in controlled and semi-synthetic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。