自动优化提示词,让AI自己搞定图文生成难题。
PromptIQ: Who Cares About Prompts? Let System Handle It -- A Component-Aware Framework for T2I Generation
- 用组件感知相似度自动检测并修正提示词结构错误
- 迭代生成直到用户满意,准确率提升显著
- 适合无提示工程经验的普通用户使用
文本到图像(T2I)模型在缺乏提示工程知识的情况下生成高质量图像仍具挑战性,往往因提示词结构不佳导致图像失真和语义错位。人类可轻松识别此类问题,但现有评估指标如CLIP无法捕捉结构不一致,暴露了当前评估方法的关键缺陷。为此,我们提出PromptIQ,一个自动化框架,通过新颖的组件感知相似度(CAS)度量来评估图像质量并优化提示词,能有效检测并惩罚结构错误。与传统方法不同,PromptIQ会持续生成并评估图像,直至用户满意,彻底消除试错式提示调优。实验表明,PromptIQ显著提升了生成质量与评估准确性,使T2I模型对无提示工程经验的用户更友好。
原文摘要 · Abstract (English)
Generating high-quality images without prompt engineering expertise remains a challenge for text-to-image (T2I) models, which often misinterpret poorly structured prompts, leading to distortions and misalignments. While humans easily recognize these flaws, metrics like CLIP fail to capture structural inconsistencies, exposing a key limitation in current evaluation methods. To address this, we introduce PromptIQ, an automated framework that refines prompts and assesses image quality using our novel Component-Aware Similarity (CAS) metric, which detects and penalizes structural errors. Unlike conventional methods, PromptIQ iteratively generates and evaluates images until the user is satisfied, eliminating trial-and-error prompt tuning. Our results show that PromptIQ significantly improves generation quality and evaluation accuracy, making T2I models more accessible for users with little to no prompt engineering expertise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。