arXiv:2511.20647cs.CV2025-11被引 2

用数学方法让同一个文字提示生成更多样化的视频。

Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization

  • 引入行列式点过程和分组相对策略优化,显式奖励多样性
  • 在多个基准上提升视频多样性,同时保持画面质量和提示一致性
  • 适用于各类扩散模型,适合视频生成研究者使用

尽管近期文本到视频(T2V)扩散模型在质量和提示对齐方面取得了显著进展,但在单个文本提示下生成多视频时往往多样性不足。本文将该问题建模为集合级策略优化问题,旨在训练一个能覆盖给定提示下多种合理结果的策略。为此,提出DPP-GRPO框架,结合行列式点过程(DPP)与分组相对策略优化(GRPO)理论,通过DPP对重复样本施加递减回报,利用GRPO提供候选集的全局反馈,使多样性成为可显式优化的目标。该框架具有即插即用、模型无关特性,在视觉外观、镜头运动和场景结构上均有效提升多样性,且不牺牲提示保真度或感知质量。我们在WAN和CogVideoX上实现该方法,在VBench、VideoScore及人工偏好评估中持续提升视频多样性。此外,我们开源代码与包含3万条多样化提示的新基准数据集,以支持后续研究。

原文摘要 · Abstract (English)

While recent text-to-video (T2V) diffusion models have achieved impressive quality and prompt alignment, they often produce low-diversity outputs when sampling multiple videos from a single text prompt. We tackle this challenge by formulating it as a set-level policy optimization problem, with the goal of training a policy that can cover the diverse range of plausible outcomes for a given prompt. To address this, we introduce DPP-GRPO, a novel framework for diverse video generation that combines Determinantal Point Processes (DPPs) and Group Relative Policy Optimization (GRPO) theories to enforce explicit reward on diverse generations. Our objective turns diversity into an explicit signal by imposing diminishing returns on redundant samples (via DPP) while supplies groupwise feedback over candidate sets (via GRPO). Our framework is plug-and-play and model-agnostic, and encourages diverse generations across visual appearance, camera motions, and scene structure without sacrificing prompt fidelity or perceptual quality. We implement our method on WAN and CogVideoX, and show that our method consistently improves video diversity on state-of-the-art benchmarks such as VBench, VideoScore, and human preference studies. Moreover, we release our code and a new benchmark dataset of 30,000 diverse prompts to support future research.

视频生成多样性扩散模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。