arXiv:2608.12290cs.CVcs.AI2026-08

用智能代理自动优化图像生成视频,减少试错,提升可控性。

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

论文配图:Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
图 1 · 摘自论文原文
  • 用大模型迭代优化提示词,自动检测语义偏差和错误。
  • 结合贝叶斯优化,高效搜索随机种子与参数组合,提升输出一致性。
  • 适合需要稳定可控视频生成的创作者和工业级应用。

现代黑箱图像到视频(I2V)模型在自动化内容创作中能力强大,但缺乏细粒度控制与可靠性,其固有的随机性导致提示词或超参数微小变化即产生显著差异,常需低效的暴力试错。为此,我们提出“智能体自优化”框架,将视频生成重构为闭环、目标导向的优化过程。第一阶段通过多模态大模型(mLLM)迭代优化提示词,结合戴维森场景图(DSG)查询确保语义一致,以及常见错误问题(CMQ)检测生成缺陷。第二阶段采用贝叶斯优化联合优化随机种子与CFG尺度,以视频-文本一致性(VTA)等质量指标为引导。人类偏好测试显示,该方法生成视频胜率高达69%,显著优于基线。本工作为提升顶尖视频生成模型的可预测性与可控性提供实用可扩展方案,推动领域从实验性探索迈向生产可用工具。

原文摘要 · Abstract (English)

Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperparameters to yield drastically different outputs often necessitating inefficient, brute-force trial-and-error processes. To address these limitations, we introduce the ``Agentic Self-Improvement" framework, which reframes video synthesis into a closed-loop, goal-directed optimization. Our framework systematically navigates the generation parameter space using a novel two-stage approach. In the first stage, an iterative prompt optimization loop uses a multimodal Large Language Model (mLLM) to refine the input prompt. This refinement implements two automated evaluations: Davidsonian Scene Graph (DSG) queries ensure semantic adherence, and Common Mistake Questions (CMQ) for artifact detection. At the second stage, we use Bayesian optimization to efficiently co-optimize stochastic seeds and CFG scales. This search is guided by a suite of quality metrics, including the novel Video-Text Adherence (VTA) score derived from the DSG and CMQ evaluations. Our framework significantly outperforms unguided search methods: in human preference studies, videos generated via our agentic approach were strongly preferred over baseline outputs, achieving win rates up to 69\%. This work provides a practical and extensible methodology for enhancing the predictability and control of state-of-the-art video generation models, moving the field beyond speculative curiosities toward reliable, production-ready tools.

视频生成智能体优化可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。