无需人工标注,用最优传输对齐奖励提升文本生成视频质量
PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards
- 用双层最优传输对齐文本与视频嵌入,生成质量与语义双重奖励
- 在VBench上超越有标注和无标注方法,在质量和语义上均领先
- 兼容反向传播与强化学习,适合多种模型优化框架
文本到视频(T2V)生成旨在合成视觉质量高、时间连贯且语义与输入文本一致的视频。基于奖励的后训练方法成为提升生成视频质量与语义对齐性的有力方向。然而,现有方法或依赖大规模人工偏好标注,或使用预训练视觉语言模型中错位的嵌入,导致可扩展性差或监督效果不佳。我们提出$ exttt{PISCES}$,一种无标注的后训练算法,通过新颖的双层最优传输(OT)对齐奖励模块解决上述问题。为使奖励信号更贴近人类判断,$ exttt{PISCES}$在分布层面和离散标记层面使用OT对齐文本与视频嵌入,实现双重目标:(i) 分布级OT对齐的质量奖励,捕捉整体视觉质量与时序一致性;(ii) 标记级OT对齐的语义奖励,强制文本与视频在时空上的对应关系。据我们所知,$ exttt{PISCES}$是首个通过OT视角改进生成后训练中无标注奖励监督的方法。在短时和长时视频生成任务上,$ exttt{PISCES}$在VBench上均优于有标注与无标注方法,质量与语义得分全面领先。人类偏好实验进一步验证其有效性。我们还证明双层OT对齐奖励模块兼容多种优化范式,包括直接反向传播与强化学习微调。
原文摘要 · Abstract (English)
Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training has emerged as a promising direction to improve the quality and semantic alignment of generated videos. However, recent methods either rely on large-scale human preference annotations or operate on misaligned embeddings from pre-trained vision-language models, leading to limited scalability or suboptimal supervision. We present $\texttt{PISCES}$, an annotation-free post-training algorithm that addresses these limitations via a novel Dual Optimal Transport (OT)-aligned Rewards module. To align reward signals with human judgment, $\texttt{PISCES}$ uses OT to bridge text and video embeddings at both distributional and discrete token levels, enabling reward supervision to fulfill two objectives: (i) a Distributional OT-aligned Quality Reward that captures overall visual quality and temporal coherence; and (ii) a Discrete Token-level OT-aligned Semantic Reward that enforces semantic, spatio-temporal correspondence between text and video tokens. To our knowledge, $\texttt{PISCES}$ is the first to improve annotation-free reward supervision in generative post-training through the lens of OT. Experiments on both short- and long-video generation show that $\texttt{PISCES}$ outperforms both annotation-based and annotation-free methods on VBench across Quality and Semantic scores, with human preference studies further validating its effectiveness. We show that the Dual OT-aligned Rewards module is compatible with multiple optimization paradigms, including direct backpropagation and reinforcement learning fine-tuning. Project page: https://roar-ai.github.io/pisces
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。