arXiv:2607.16947cs.CV2026-07中稿 · ACM MM 2026被引 2

解决文本生成视频中物理真实与语义一致的矛盾

When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation

论文配图:When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation
图 1 · 摘自论文原文
  • 通过融合物理与语义信号,动态调整偏好对权重
  • 在VideoPhy-2上物理真实性提升2倍,语义一致性不下降
  • 无需额外模型或损失项,可直接集成到标准DPO框架

文本到视频生成模型虽具强视觉真实感,但提升物理合理性常以牺牲输入文本的语义一致性为代价。这一矛盾源于传统物理偏好依赖视频动态对比,却未考量视频是否忠实呈现提示描述的场景,导致物理-语义冲突成为系统性问题。本文将其建模为约束偏好优化问题,提出物理与语义直接偏好优化(PSDPO),根据物理与语义信号的一致性动态调节每对偏好贡献。梯度级分析表明,PSDPO将冲突对引起的语义漂移控制在可调节残差内,并进一步提出分阶段优化协议,可证明减少累积漂移。方法完全基于标准DPO框架,无需辅助模型或额外损失项。实验显示,PSDPO在VideoPhy-2上物理可塑性提升达2倍,同时在VBench上保持强语义一致性,优于现有基于偏好的方法。

原文摘要 · Abstract (English)

Text-to-video (T2V) generation models have achieved strong visual realism, but improving physical plausibility can come at the cost of semantic consistency with the input text. This tension arises because physical preference is typically determined by comparing dynamics between two videos, without accounting for whether either video faithfully depicts the scene specified by the prompt, making physical-semantic conflict a systematic tendency under this supervision paradigm. We formulate this challenge as a constrained preference optimization problem and propose Physical and Semantic Direct Preference Optimization (PSDPO), which modulates each preference pair's contribution based on the agreement between its physical and semantic signals. A gradient-level analysis shows that PSDPO bounds the semantic drift from conflicting pairs to a controllable residual, and further motivates a staged optimization protocol that provably reduces cumulative drift. The resulting method operates entirely within the standard DPO framework, requiring no auxiliary models or additional loss terms. Experiments show that PSDPO improves physical plausibility by up to $2\times$ over the baseline on VideoPhy-2, while maintaining strong semantic consistency on VBench, achieving a more reliable balance than existing preference-based methods.

文本生成视频偏好优化物理合理性语义一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。