arXiv:2605.28230cs.CV2026-05被引 2

让视频生成模型自检物理合理性,无需训练即可提升真实感

Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation

论文配图:Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation
图 1 · 摘自论文原文
  • 通过潜空间扰动捕捉模型自身运动残差,作为物理合理性评分信号
  • 在TurboWan2.2上使Physics-IQ提升16.5%,VideoPhy2-hard提升20.6%
  • 适合追求真实物理行为的视频生成研究者和应用开发者

现代视频生成模型虽视觉效果出色,但常违背基本物理规律。本文提出Proprio,一种无需训练的框架,使冻结的视频生成器能自主评估并改进输出的物理合理性。受本体感知启发,Proprio将模型在受控潜空间扰动下的运动残差视为自我评分信号:更符合模型学习动态的样本产生更小、更稳定的残差。该信号在时间步与扰动间聚合,并通过动态时空掩码聚焦于运动相关区域,用于最佳样本选择、梯度自优化或二者结合。在文本到视频与图像到视频基准测试中,Proprio持续提升物理合理性,优于基于视觉语言模型的评分与外部世界模型基线。使用TurboWan2.2时,Physics-IQ从32.2升至37.5(+16.5%),VideoPhy2-hard物理常识任务从45.6升至55.0(+20.6%)。人工评估显示,在约三分之二的对比中,评判者更偏好Proprio筛选或优化后的视频。结果表明,冻结的视频生成器内含可利用的内部信号,可用于自主评估与提升其输出的物理合理性。

原文摘要 · Abstract (English)

Modern video generative models produce visually impressive results, yet frequently violate basic physical principles. We propose Proprio, a training-free framework that enables a frozen video generator to assess and improve the physical plausibility of its own outputs. Inspired by proprioception, the biological sense of one's own movement, Proprio treats the model's flow residual under controlled latent perturbations as a self-scoring signal. Samples that are better explained by the generator's learned dynamics induce smaller and more stable residuals. We aggregate this signal across timesteps and perturbations, focus it on motion-relevant regions with a dynamic spatiotemporal mask, and use it for best-of-N search, gradient-based self-refinement, or both. Across text-to-video and image-to-video benchmarks, Proprio consistently improves physical plausibility, outperforming VLM-based scoring, and external world-model baselines in several settings. With TurboWan2.2, Proprio improves Physics-IQ from 32.2 to 37.5 (+16.5%) and VideoPhy2-hard physical commonsense from 45.6 to 55.0 (+20.6%). Human evaluation further shows that raters prefer Proprio-selected or refined videos for physical plausibility in roughly two-thirds of comparisons. These results suggest that frozen video generators contain actionable internal signals for evaluating and improving the physical plausibility of their own outputs.

视频生成物理合理性自评分零训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。