arXiv:2603.01509cs.CVcs.AI2026-03

无需训练即可提升文本生成视频质量,通过优化提示实现更精准连贯的视频输出。

Retrieval, Refinement, and Ranking for Text-to-Video Generation via Prompt Optimization and Test-Time Scaling

  • 基于检索增强生成的提示优化框架,不需微调主模型。
  • 提升视频静态清晰度与动态连贯性,生成效果更贴近人类偏好。
  • 适合希望提升生成质量但无资源训练模型的研究者与开发者。

尽管大规模数据集推动了文本到视频(T2V)生成模型的显著进展,这些模型仍对输入提示高度敏感,表明提示设计对生成质量至关重要。当前改进视频输出的方法往往存在局限:要么依赖复杂的后期编辑模型,易引入伪影;要么需要昂贵的主生成器微调,严重限制可扩展性和可访问性。本文提出3R,一种基于检索增强生成(RAG)的提示优化框架。3R利用当前最先进的T2V扩散模型和视觉语言模型,可适配任意T2V模型且无需任何模型训练。该框架采用三项核心策略:基于RAG的修饰词提取以增强上下文关联性,基于扩散模型的偏好优化以对齐人类偏好,以及时间帧插值以生成时序一致的视觉内容。三者协同实现更准确、高效且语境一致的文本到视频生成。实验结果表明,3R能有效提升生成视频的静态保真度与动态连贯性,凸显优化用户提示的重要性。

原文摘要 · Abstract (English)

While large-scale datasets have driven significant progress in Text-to-Video (T2V) generative models, these models remain highly sensitive to input prompts, demonstrating that prompt design is critical to generation quality. Current methods for improving video output often fall short: they either depend on complex, post-editing models, risking the introduction of artifacts, or require expensive fine-tuning of the core generator, which severely limits both scalability and accessibility. In this work, we introduce 3R, a novel RAG based prompt optimization framework. 3R utilizes the power of current state-of-the-art T2V diffusion model and vision language model. It can be used with any T2V model without any kind of model training. The framework leverages three key strategies: RAG-based modifiers extraction for enriched contextual grounding, diffusion-based Preference Optimization for aligning outputs with human preferences, and temporal frame interpolation for producing temporally consistent visual contents. Together, these components enable more accurate, efficient, and contextually aligned text-to-video generation. Experimental results demonstrate the efficacy of 3R in enhancing the static fidelity and dynamic coherence of generated videos, underscoring the importance of optimizing user prompts.

文本生成视频提示优化扩散模型RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。