arXiv:2603.14294cs.CVcs.AI2026-03

发现扩散模型中间层能区分物理合理与不合理视频,无需微调即可提升生成质量。

Seeking Physics in Diffusion Noise

  • 在高噪声下仍能通过中间特征分离物理合理与不合理视频
  • 两种推理机制使生成视频物理一致性显著提升,提速37%
  • 适用于无需微调的视频生成优化,适合追求真实感的场景

视频扩散模型是否编码了预测物理合理性信号?我们探测了预训练扩散变换器(DiTs)中间去噪表征,发现即使在高噪声水平下,物理合理与不合理视频在中层特征空间中仍可部分分离。源内和感知质量控制表明该信号并非完全由生成器身份或通用视觉质量解释。我们将此信号提炼为轻量级、针对特定主干网络的物理验证器,基于冻结特征训练,并应用于两种互补的推理时机制:渐进轨迹选择(在中间检查点评分并早期剪枝弱候选)与奖励梯度引导(仅反向传播前几层以引导剩余轨迹)。在 PhyGenBench 与 Physics-IQ 数据集上对 CogVideoX-2B/5B 与 Wan 2.1-14B 的实验表明,渐进选择在 CogVideoX-2B 上达到基于验证器的 Best-of-4 效果,同时将实际推理时间减少 37%;奖励梯度引导在 CogVideoX-5B 上显著提升物理一致性,全程无需微调视频生成器。

原文摘要 · Abstract (English)

Do video diffusion models encode signals predictive of physical plausibility? We probe intermediate denoising representations of pretrained Diffusion Transformers (DiTs) and find that physically plausible and implausible videos are partially separable in mid-layer feature space, even at high noise levels. Within-source and perceptual-quality controls suggest that this signal is not fully explained by generator identity or generic visual quality. We distill the signal into a lightweight, backbone-specific physics verifier trained on frozen features and use it in two complementary inference-time mechanisms under a fixed multi-trajectory budget: progressive trajectory selection, which scores trajectories at intermediate checkpoints and prunes weak candidates early, and reward-gradient guidance, which steers surviving trajectories by backpropagating through only the first few DiT blocks. Experiments on PhyGenBench and Physics-IQ across CogVideoX-2B/5B and Wan 2.1-14B show that progressive selection matches verifier-based Best-of-4 on CogVideoX-2B while reducing wall-clock inference time by 37%, whereas reward-gradient guidance substantially improves physical consistency on CogVideoX-5B, all without fine-tuning the video generator.

视频生成扩散模型物理一致性推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。