让视频生成更符合物理规律,通过解耦语义与物理实现精准对比学习。
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
- 解耦语义与物理,设计宏观和微观双尺度对比信号
- 在VideoPhy上物理常识得分提升16.7%(相比基线)
- 无需额外训练时间,适合希望提升物理合理性视频模型的用户
流匹配视频生成器虽能生成时间连贯、高保真的视频,但常违背基本物理规律,因其重建目标惩罚每帧偏差,却无法区分物理合理与不合理动态。对比流匹配通过推远不同条件下的速度场轨迹提供理论解决方案,但我们发现文本条件视频设置中存在语义-物理纠缠的根本障碍:自然语言提示将场景内容与物理行为耦合,导致朴素负样本采样所选条件的速度场与正样本高度重叠,使对比梯度直接对抗流匹配目标。我们形式化了这一梯度冲突,推导出精确对齐条件,揭示对比学习在何种情况下有益或有害。基于此分析,提出轻量级后训练框架DiReCT(Disentangled Regularization of Contrastive Trajectories),将对比信号分解为两个互补尺度:宏观对比项从语义迥异区域提取不重叠负样本,实现无干扰的全局轨迹分离;微观对比项构建与正样本共享完整场景语义但仅在一个由大语言模型扰动的物理维度上不同的困难负样本,覆盖运动学、受力、材料、相互作用及量级等。速度空间分布正则化防止预训练视觉质量的灾难性遗忘。应用于Wan 2.1-1.3B时,该方法在VideoPhy上相比基线和SFT分别提升物理常识得分16.7%和11.3%,且不增加训练时间。
原文摘要 · Abstract (English)
Flow-matching video generators produce temporally coherent, high-fidelity outputs yet routinely violate elementary physics because their reconstruction objectives penalize per-frame deviations without distinguishing physically consistent dynamics from impossible ones. Contrastive flow matching offers a principled remedy by pushing apart velocity-field trajectories of differing conditions, but we identify a fundamental obstacle in the text-conditioned video setting: semantic-physics entanglement. Because natural-language prompts couple scene content with physical behavior, naive negative sampling draws conditions whose velocity fields largely overlap with the positive sample's, causing the contrastive gradient to directly oppose the flow-matching objective. We formalize this gradient conflict, deriving a precise alignment condition that reveals when contrastive learning helps versus harms training. Guided by this analysis, we introduce DiReCT (Disentangled Regularization of Contrastive Trajectories), a lightweight post-training framework that decomposes the contrastive signal into two complementary scales: a macro-contrastive term that draws partition-exclusive negatives from semantically distant regions for interference-free global trajectory separation, and a micro-contrastive term that constructs hard negatives sharing full scene semantics with the positive sample but differing along a single, LLM-perturbed axis of physical behavior; spanning kinematics, forces, materials, interactions, and magnitudes. A velocity-space distributional regularizer helps to prevent catastrophic forgetting of pretrained visual quality. When applied to Wan 2.1-1.3B, our method improves the physical commonsense score on VideoPhy by 16.7% and 11.3% compared to the baseline and SFT, respectively, without increasing training time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。