用多个智能体自动优化文本生成图像的提示词,提升质量并减少修改次数。
PromptSculptor: Multi-Agent Based Text-to-Image Prompt Optimization
- 四类专业智能体协作,逐步细化模糊提示词。
- 相比手动调整,生成图像质量显著提升,迭代次数减少。
- 可适配多种文生图模型,适合工业级应用开发。
生成式AI的快速发展使文生图模型等强大工具得以普及。然而,生成高质量图像仍需用户精心设计包含场景、风格和上下文的详细提示词,常需多轮反复调整。本文提出PromptSculptor,一种基于多智能体的新型框架,自动化迭代优化提示词过程。系统将任务分解为四类专业化智能体协同工作,将初始简短模糊的用户输入转化为完整精细的提示。通过链式思维推理,框架能有效推断隐藏上下文,并丰富场景与背景细节。为实现持续优化,自评估智能体确保修改后提示与原始输入一致,反馈调优智能体则融合用户反馈进行进一步改进。实验表明,PromptSculptor显著提升生成图像质量,大幅减少用户达到满意结果所需的迭代次数。其模型无关的设计特性使其可无缝集成于各类Text-to-Image(T2I)模型,为工业应用铺平道路。
原文摘要 · Abstract (English)
The rapid advancement of generative AI has democratized access to powerful tools such as Text-to-Image models. However, to generate high-quality images, users must still craft detailed prompts specifying scene, style, and context-often through multiple rounds of refinement. We propose PromptSculptor, a novel multi-agent framework that automates this iterative prompt optimization process. Our system decomposes the task into four specialized agents that work collaboratively to transform a short, vague user prompt into a comprehensive, refined prompt. By leveraging Chain-of-Thought reasoning, our framework effectively infers hidden context and enriches scene and background details. To iteratively refine the prompt, a self-evaluation agent aligns the modified prompt with the original input, while a feedback-tuning agent incorporates user feedback for further refinement. Experimental results demonstrate that PromptSculptor significantly enhances output quality and reduces the number of iterations needed for user satisfaction. Moreover, its model-agnostic design allows seamless integration with various T2I models, paving the way for industrial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。