arXiv:2507.20536cs.CVcs.AI2025-07ICCV被引 31

无需训练的多智能体系统,自动优化文本生成图像的提示词与质量。

T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation

论文配图:T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation
图 1 · 摘自论文原文
  • 三个智能体协作解析提示、选模型、评估生成结果。
  • 在GenAI-Bench上性能媲美商用模型,成本仅为1/6。
  • 支持全自动或人工介入,提升图像与文本一致性。

文本到图像(T2I)生成模型虽革新了内容创作,但对提示词敏感,常需反复修改且缺乏明确反馈。现有技术如自动提示工程、控制嵌入、去噪和多轮生成存在可控性差或需额外训练的问题。为此,我们提出T2I-Copilot——一个无需训练的多智能体系统,通过(多模态)大语言模型协作,实现提示词标准化、模型选择与迭代优化。系统包含三部分:(1) 输入解释器,解析提示并消除歧义;(2) 生成引擎,从不同T2I模型中选型并组织图文提示;(3) 质量评估器,评估美学与图文对齐度并提供反馈。T2I-Copilot可全自主运行,也支持人工干预。在GenAI-Bench测试中,使用开源模型,其VQA得分接近RecraftV3与Imagen 3,优于FLUX1.1-pro 6.17%(仅为其16.59%成本),超越FLUX.1-dev与SD 3.5 Large分别达9.11%与6.36%。

原文摘要 · Abstract (English)

Text-to-Image (T2I) generative models have revolutionized content creation but remain highly sensitive to prompt phrasing, often requiring users to repeatedly refine prompts multiple times without clear feedback. While techniques such as automatic prompt engineering, controlled text embeddings, denoising, and multi-turn generation mitigate these issues, they offer limited controllability, or often necessitate additional training, restricting the generalization abilities. Thus, we introduce T2I-Copilot, a training-free multi-agent system that leverages collaboration between (Multimodal) Large Language Models to automate prompt phrasing, model selection, and iterative refinement. This approach significantly simplifies prompt engineering while enhancing generation quality and text-image alignment compared to direct generation. Specifically, T2I-Copilot consists of three agents: (1) Input Interpreter, which parses the input prompt, resolves ambiguities, and generates a standardized report; (2) Generation Engine, which selects the appropriate model from different types of T2I models and organizes visual and textual prompts to initiate generation; and (3) Quality Evaluator, which assesses aesthetic quality and text-image alignment, providing scores and feedback for potential regeneration. T2I-Copilot can operate fully autonomously while also supporting human-in-the-loop intervention for fine-grained control. On GenAI-Bench, using open-source generation models, T2I-Copilot achieves a VQA score comparable to commercial models RecraftV3 and Imagen 3, surpasses FLUX1.1-pro by 6.17% at only 16.59% of its cost, and outperforms FLUX.1-dev and SD 3.5 Large by 9.11% and 6.36%. Code will be released at: https://github.com/SHI-Labs/T2I-Copilot.

文本生成图像多智能体提示优化生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。