arXiv:2607.28058cs.CV2026-07

通过重建误差隐式优化视频生成的时序一致性,解决运动模糊等缺陷。

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion

论文配图:Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion
图 1 · 摘自论文原文
  • 从去噪过程提取隐式偏好信号,无需人工标注。
  • 聚焦高误差帧段进行优化,显著减少运动崩塌等问题。
  • 适用于后训练阶段提升扩散模型生成视频的真实感。

基于扩散模型的文本到视频生成在偏好对齐方面取得进展,尤其是通过直接偏好优化(DPO)提升了视觉质量。然而,运动崩塌、物体闪烁和色彩过饱和等时序稀疏伪影仍是感知真实性的主要障碍。现有方法受限于两个关键问题:(1) 偏好归属瓶颈——离线人工标注成本高且难以捕捉学习动态,而在线奖励信号虽具回滚感知但常不稳定且有偏差;(2) 时序信用误分配——均匀施加的监督无法有效定位伪影出现的短暂片段。为此,我们提出集中式隐式偏好优化(cIPO),一种面向视频扩散模型的后训练框架。cIPO直接从去噪过程中推导隐式偏好信号:给定真实视频,模型添加前向噪声并迭代去噪重建,将原视频视为优选样本,重建结果视为次优样本。该设定在无需人工标注或外部奖励模型的情况下捕捉推理时误差。此外,原始与重建视频之间的帧级差异揭示了失败发生的时间点。cIPO利用此信息计算时序重建误差,并将优化集中在高误差片段,实现对易出错区域的更精准修正。大量实验表明,cIPO在多个数据集上持续提升视频真实性和时序连贯性,验证了隐式偏好与时序集中优化的有效性与高效性。

原文摘要 · Abstract (English)

Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly improved visual quality. However, temporally sparse artifacts such as motion collapse, object flickering, and color oversaturation remain a major barrier to perceptual realism. Existing methods struggle with these issues due to two key limitations: (1) the preference attribution bottleneck, where offline human annotations are costly and fail to accurately capture learning dynamics, while online reward signals are rollout-aware but often unstable and biased; and (2) temporal credit misallocation, where uniformly applied supervision cannot effectively target the brief segments in which artifacts occur. To address these challenges, we propose concentrated Implicit Preference Optimization (cIPO), a post-training framework for video diffusion models. cIPO derives implicit preference signals directly from the denoising process: given a real video, the model adds forward noise and reconstructs it via iterative denoising, treating the original as the preferred sample and the reconstruction as the dispreferred one. This formulation captures inference-time errors without requiring human annotations or external reward models. Moreover, frame-level discrepancies between original and reconstructed videos reveal when failures occur. cIPO leverages this by computing temporal reconstruction errors and concentrating optimization on high-error segments, enabling more precise correction of failure-prone regions. Extensive experiments demonstrate that cIPO consistently enhances video authenticity and temporal coherence across multiple datasets, highlighting the effectiveness and efficiency of implicit preference with temporally concentrated optimization.

视频生成扩散模型偏好优化时序一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。