arXiv:2509.10704cs.AIcs.CV2025-09被引 14

让AI自动优化图像生成,无需人工反复改提示词。

Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration

  • 用多模态大模型当评委,自动发现图像问题并给出可解释修改建议。
  • 通过对比迭代生成的图像,逐步进化出更符合用户意图的高质量结果。
  • 适合想省去繁琐调参、追求高质量图像生成的创作者和研究者。

文本到图像(T2I)模型虽具巨大创意潜力,但高度依赖人工干预,常需对模糊提示进行手动、反复的提示工程。本文提出Maestro,一种自演化图像生成系统,使T2I模型仅凭初始提示即可自主改进生成图像。其核心创新包括:1)自评机制,由专用多模态大模型(MLLM)代理充当“批评者”,识别图像缺陷、补全描述不足,并提供可解释的编辑信号,由“验证者”代理整合,同时保留用户意图;2)自演化机制,利用MLLM作为裁判,进行图像间头对头比较,剔除低质量图像,进化出更契合用户意图的创造性提示候选。在复杂T2I任务上使用黑盒模型的大量实验表明,Maestro显著优于初始提示和现有自动化方法,且性能随更先进的MLLM组件提升。该工作为自改进T2I生成提供了鲁棒、可解释且高效的新路径。

原文摘要 · Abstract (English)

Text-to-image (T2I) models, while offering immense creative potential, are highly reliant on human intervention, posing significant usability challenges that often necessitate manual, iterative prompt engineering over often underspecified prompts. This paper introduces Maestro, a novel self-evolving image generation system that enables T2I models to autonomously self-improve generated images through iterative evolution of prompts, using only an initial prompt. Maestro incorporates two key innovations: 1) self-critique, where specialized multimodal LLM (MLLM) agents act as 'critics' to identify weaknesses in generated images, correct for under-specification, and provide interpretable edit signals, which are then integrated by a 'verifier' agent while preserving user intent; and 2) self-evolution, utilizing MLLM-as-a-judge for head-to-head comparisons between iteratively generated images, eschewing problematic images, and evolving creative prompt candidates that align with user intents. Extensive experiments on complex T2I tasks using black-box models demonstrate that Maestro significantly improves image quality over initial prompts and state-of-the-art automated methods, with effectiveness scaling with more advanced MLLM components. This work presents a robust, interpretable, and effective pathway towards self-improving T2I generation.

图像生成自进化多模态提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。