arXiv:2412.15156cs.CVcs.CL2024-12ICCV被引 21

用AI自动优化视频生成提示词,让输出更符合用户偏好。

Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM

论文配图:Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
图 1 · 摘自论文原文
  • 用大模型分两阶段优化提示词,生成更贴合视频模型的描述。
  • 在多个视频生成模型上测试,显著提升输出质量与一致性。
  • 适合想高效生成高质量视频的创作者和研究者使用。

文本到视频模型通过高质量文本-视频对的优化取得了显著进展,其中文本提示在决定输出视频质量方面起着关键作用。然而,实现理想输出通常需要多次修改和迭代推理来完善用户提供的提示。现有的自动提示优化方法在应用于文本到视频扩散模型时面临模态不一致、成本差异和模型无感知等挑战。为此,我们提出一种基于大语言模型的提示词自适应框架——Prompt-A-Video,该框架能生成以视频为中心、无需人工干预且与用户偏好对齐的提示词,适配特定视频扩散模型。我们的方法包含精心设计的两阶段优化与对齐系统:首先通过奖励引导的提示演化流程自动构建最优提示池,并用于大语言模型的监督微调(SFT);随后采用多维度奖励生成成对数据,再利用直接偏好优化(DPO)算法进一步实现偏好对齐。通过大量实验与对比分析,我们在多种生成模型上验证了Prompt-A-Video的有效性,展示了其推动视频生成边界的能力。

原文摘要 · Abstract (English)

Text-to-video models have made remarkable advancements through optimization on high-quality text-video pairs, where the textual prompts play a pivotal role in determining quality of output videos. However, achieving the desired output often entails multiple revisions and iterative inference to refine user-provided prompts. Current automatic methods for refining prompts encounter challenges such as Modality-Inconsistency, Cost-Discrepancy, and Model-Unaware when applied to text-to-video diffusion models. To address these problem, we introduce an LLM-based prompt adaptation framework, termed as Prompt-A-Video, which excels in crafting Video-Centric, Labor-Free and Preference-Aligned prompts tailored to specific video diffusion model. Our approach involves a meticulously crafted two-stage optimization and alignment system. Initially, we conduct a reward-guided prompt evolution pipeline to automatically create optimal prompts pool and leverage them for supervised fine-tuning (SFT) of the LLM. Then multi-dimensional rewards are employed to generate pairwise data for the SFT model, followed by the direct preference optimization (DPO) algorithm to further facilitate preference alignment. Through extensive experimentation and comparative analyses, we validate the effectiveness of Prompt-A-Video across diverse generation models, highlighting its potential to push the boundaries of video generation.

视频生成提示优化大模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。