arXiv:2509.12446cs.MAcs.AI2025-09EMNLP被引 34

用多个智能体自动优化文本生成图像的提示词,提升质量并减少修改次数。

PromptSculptor: Multi-Agent Based Text-to-Image Prompt Optimization

  • 四类专业智能体协作,逐步细化模糊提示词。
  • 相比手动调整,生成图像质量显著提升,迭代次数减少。
  • 可适配多种文生图模型,适合工业级应用开发。

生成式AI的快速发展使文生图模型等强大工具得以普及。然而,生成高质量图像仍需用户精心设计包含场景、风格和上下文的详细提示词,常需多轮反复调整。本文提出PromptSculptor,一种基于多智能体的新型框架,自动化迭代优化提示词过程。系统将任务分解为四类专业化智能体协同工作,将初始简短模糊的用户输入转化为完整精细的提示。通过链式思维推理,框架能有效推断隐藏上下文,并丰富场景与背景细节。为实现持续优化,自评估智能体确保修改后提示与原始输入一致,反馈调优智能体则融合用户反馈进行进一步改进。实验表明,PromptSculptor显著提升生成图像质量,大幅减少用户达到满意结果所需的迭代次数。其模型无关的设计特性使其可无缝集成于各类Text-to-Image(T2I)模型,为工业应用铺平道路。

原文摘要 · Abstract (English)

The rapid advancement of generative AI has democratized access to powerful tools such as Text-to-Image models. However, to generate high-quality images, users must still craft detailed prompts specifying scene, style, and context-often through multiple rounds of refinement. We propose PromptSculptor, a novel multi-agent framework that automates this iterative prompt optimization process. Our system decomposes the task into four specialized agents that work collaboratively to transform a short, vague user prompt into a comprehensive, refined prompt. By leveraging Chain-of-Thought reasoning, our framework effectively infers hidden context and enriches scene and background details. To iteratively refine the prompt, a self-evaluation agent aligns the modified prompt with the original input, while a feedback-tuning agent incorporates user feedback for further refinement. Experimental results demonstrate that PromptSculptor significantly enhances output quality and reduces the number of iterations needed for user satisfaction. Moreover, its model-agnostic design allows seamless integration with various T2I models, paving the way for industrial applications.

文生图提示优化多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。