arXiv:2607.04653cs.CV2026-07被引 2

通过角色感知与模态解耦,提升视频生成的物理一致性

Enhancing Video Physical Consistency via Role-aware Joint Training and Modality-decoupled Denoising

论文配图:Enhancing Video Physical Consistency via Role-aware Joint Training and Modality-decoupled Denoising
图 1 · 摘自论文原文
  • 按实体角色分组,分别建模运动规律
  • 模态解耦去噪使辅助信息仅作软约束
  • 适合关注视频动态真实性的生成研究者

现代视频扩散模型在视觉保真度上表现优异,但长期物理一致性仍面临挑战。传统基于像素重建的目标主要关注外观细节,难以捕捉场景底层动态。近期方法引入光流等辅助模态,通过联合训练融入物理先验,但存在三方面缺陷:(1) 未区分不同实体类型的运动模式;(2) 视觉与辅助模态联合建模导致容量冲突,削弱预训练视觉先验;(3) 辅助模态推理中误差累积。为此,我们提出VPT微调框架,引入角色感知信号,将实体分为主体、受控物体、被动物体和背景,实现更清晰的角色化建模。进一步提出模态解耦去噪策略,为视觉与辅助通道分配独立噪声水平,并结合损失权重衰减机制,使辅助模态作为软约束而非强依赖,缓解推理中的递归预测误差。还引入跨步自引导以增强物理动态。实验表明,VPT在VideoPhy基准上相较Wan2.1-T2V-1.3B提升SA 39.4%、PC 17.9%,VideoPhy-2上也取得一致改进。

原文摘要 · Abstract (English)

While modern video diffusion models excel in visual fidelity, maintaining long-range physical consistency remains a formidable challenge. Conventional pixel-reconstruction objectives mainly focus on appearance details and often fail to capture the underlying dynamics of a scene. To mitigate this, recent efforts have integrated auxiliary modalities (e.g., optical flow) to introduce physics priors via joint training with video appearance. However, these methods have three main limitations: (1) they do not distinguish the different motion patterns of different entity types; (2) joint modeling of visual and auxiliary modalities can cause capacity conflicts and weaken the pretrained visual prior; and (3) auxiliary modalities may accumulate errors during inference. To address these issues, we propose \textbf{VPT}, a fine-tuning framework for improving physical consistency in video diffusion models. VPT introduces a role-aware signal that groups entities into agents, controlled objects, passive objects, and background, so that different physical roles can be modeled more clearly. We further propose a modality-decoupled denoising strategy, where the visual and auxiliary channels are assigned independent noise levels. Together with a loss-weight decay strategy, this design makes auxiliary modalities serve as soft constraints rather than strong dependencies, mitigating recursive prediction errors during inference. We also introduce cross-step auto-guidance to further strengthen physical dynamics. Experiments show that VPT improves physical consistency while preserving visual quality, achieving relative gains of 39.4\% in SA and 17.9\% in PC on VideoPhy benchmark over Wan2.1-T2V-1.3B, and consistent improvements on VideoPhy-2 benchmark. The project page is available at https://tom-zgt.github.io/VPT.

视频生成物理一致性扩散模型模态解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。