arXiv:2502.02088cs.CVcs.AI2025-02被引 15

通过双迭代优化提升文本生成视频的真人偏好对齐与质量

Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation

  • 交替优化奖励模型与生成模型,实现双向增强
  • 2B参数模型经优化后超越未优化的5B模型
  • 无需人工标注,自动提升画面一致性与流畅度

近期视频生成进展已实现由可扩展扩散变压器驱动的真实视频生成。然而,这些方法通常难以产出符合用户真实需求和偏好的满意结果。本文提出双迭代优化(Dual-IPO),一种迭代范式,依次优化奖励模型与视频生成模型,以提升合成质量与人类偏好对齐。奖励模型通过思维链引导推理、基于投票的自一致性及偏好确定性估计,确保可靠且稳健的奖励信号。基于此,我们利用奖励模型反馈指导视频基础模型优化,从而提升主体一致性、运动流畅性与美学质量等。奖励模型与生成模型在多轮迭代中相互促进,持续改进,无需繁琐的人工偏好标注。大量实验表明,所提Dual-IPO能有效且一致地提升不同架构与规模基模型的视频生成质量,甚至使仅20亿参数的模型超越50亿参数的基线模型。分析实验与消融研究验证了系统设计合理性及各组件有效性。

原文摘要 · Abstract (English)

Recent advances in video generation have enabled thrilling experiences in producing realistic videos driven by scalable diffusion transformers. However, they usually fail to produce satisfactory outputs that are aligned to users' authentic demands and preferences. In this work, we introduce Dual-Iterative Optimization (Dual-IPO), an iterative paradigm that sequentially optimizes both the reward model and the video generation model for improved synthesis quality and human preference alignment. For the reward model, our framework ensures reliable and robust reward signals via CoT-guided reasoning, voting-based self-consistency, and preference certainty estimation. Given this, we optimize video foundation models with guidance of signals from reward model's feedback, thus improving the synthesis quality in subject consistency, motion smoothness and aesthetic quality, etc. The reward model and video generation model complement each other and are progressively improved in the multi-round iteration, without requiring tediously manual preference annotations. Comprehensive experiments demonstrate that the proposed Dual-IPO can effectively and consistently improve the video generation quality of base model with various architectures and sizes, even help a model with only 2B parameters surpass a 5B one. Moreover, our analysis experiments and ablation studies identify the rational of our systematic design and the efficacy of each component.

文本生成视频偏好优化扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。