优化文本提示,让AI视频生成更准确
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation

- 用检索增强方法改进用户输入的提示词
- 提升视频静态画面和动态效果质量
- 适合想提升AI视频生成效果的创作者
文本到视频(T2V)生成模型虽取得显著进展,但其对输入提示的高度敏感性凸显了提示设计的关键作用。现有研究多依赖大语言模型(LLM)将用户提示与训练数据分布对齐,却缺乏对提示词汇与句法结构的精细引导。为此,本文提出一种检索增强提示优化框架RAPO。该框架通过双分支优化机制:第一分支从学习的关联图中提取多样修饰语,结合微调后的LLM使提示更贴近训练数据格式;第二分支则基于预训练LLM和明确指令集重写原始提示。大量实验表明,RAPO能有效提升生成视频在静态与动态维度的表现,验证了提示优化对用户输入提示的重要性。
原文摘要 · Abstract (English)
The evolution of Text-to-video (T2V) generative models, trained on large-scale datasets, has been marked by significant progress. However, the sensitivity of T2V generative models to input prompts highlights the critical role of prompt design in influencing generative outcomes. Prior research has predominantly relied on Large Language Models (LLMs) to align user-provided prompts with the distribution of training prompts, albeit without tailored guidance encompassing prompt vocabulary and sentence structure nuances. To this end, we introduce RAPO, a novel Retrieval-Augmented Prompt Optimization framework. In order to address potential inaccuracies and ambiguous details generated by LLM-generated prompts. RAPO refines the naive prompts through dual optimization branches, selecting the superior prompt for T2V generation. The first branch augments user prompts with diverse modifiers extracted from a learned relational graph, refining them to align with the format of training prompts via a fine-tuned LLM. Conversely, the second branch rewrites the naive prompt using a pre-trained LLM following a well-defined instruction set. Extensive experiments demonstrate that RAPO can effectively enhance both the static and dynamic dimensions of generated videos, demonstrating the significance of prompt optimization for user-provided prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。