VISTA让视频生成自动优化,通过反复改写提示词提升质量。
VISTA: A Test-Time Self-Improving Video Generation Agent
- 构建多智能体系统,迭代优化视频生成的提示词。
- 在多场景测试中,视频胜率最高达60%,人类偏好率达66.4%。
- 适合需要高质量、高对齐度视频生成的研究与应用者。
尽管文本到视频合成技术快速发展,生成视频的质量仍严重依赖用户提示的精确性。现有测试时优化方法在其他领域表现良好,但在视频这一多维度任务中效果有限。本文提出VISTA(Video Iterative Self-improvement Agent),一种新型多智能体系统,通过迭代循环自主改进视频生成。VISTA首先将用户意图分解为结构化的时间计划;生成后,通过稳健的两两锦标赛选出最佳视频;该胜出视频由视觉、音频和语境三类专业智能体分别评估;最后,推理智能体整合反馈,自我反思并改写提示词进入下一轮生成。在单场景与多场景视频生成实验中,相比先前方法,VISTA表现出一致的质量提升,对齐用户意图能力显著增强,在与最先进基线的对比中达到最高60%的胜率。人类评估结果显示,66.4%的情况下更偏爱VISTA生成结果。
原文摘要 · Abstract (English)
Despite rapid advances in text-to-video synthesis, generated video quality remains critically dependent on precise user prompts. Existing test-time optimization methods, successful in other domains, struggle with the multi-faceted nature of video. In this work, we introduce VISTA (Video Iterative Self-improvemenT Agent), a novel multi-agent system that autonomously improves video generation through refining prompts in an iterative loop. VISTA first decomposes a user idea into a structured temporal plan. After generation, the best video is identified through a robust pairwise tournament. This winning video is then critiqued by a trio of specialized agents focusing on visual, audio, and contextual fidelity. Finally, a reasoning agent synthesizes this feedback to introspectively rewrite and enhance the prompt for the next generation cycle. Experiments on single- and multi-scene video generation scenarios show that while prior methods yield inconsistent gains, VISTA consistently improves video quality and alignment with user intent, achieving up to 60% pairwise win rate against state-of-the-art baselines. Human evaluators concur, preferring VISTA outputs in 66.4% of comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。