通过视觉提问帮助用户精准生成图像,提升意图对齐效率。
Adaptive Prompt Elicitation for Text-to-Image Generation
- 用语言模型先验将用户意图转为可解释的特征需求
- 自适应生成视觉问题,用户只需看图选答而非打字
- 128人实测显示满意度提升19.8%,无需额外负担
文本到图像生成中,用户常因输入模糊或不熟悉模型特性而难以准确表达意图。本文提出自适应提示获取(APE),通过信息论框架实现交互式意图推断。APE将隐含用户意图表示为可解释的特征要求,利用语言模型先验自适应生成视觉查询,并将用户反馈整合为有效提示。在IDEA-Bench和DesignBench上的评估显示,该方法显著提升意图对齐效果并提高效率。128名参与者在自定义任务中的用户研究表明,感知对齐度提升19.8%,且未增加工作量。本工作为提示交互提供了一种系统性补充,增强了文本到图像模型的人机协同能力。
原文摘要 · Abstract (English)
Aligning text-to-image generation with user intent remains challenging, as users frequently provide ambiguous inputs and struggle with model idiosyncrasies. We propose Adaptive Prompt Elicitation (APE), a technique that adaptively poses visual queries to help users refine prompts without extensive writing. Our technical contribution is a formulation of interactive intent inference under an information-theoretic framework. APE represents latent user intent as interpretable feature requirements using language model priors, adaptively generates visual queries, and compiles elicited requirements into effective prompts. Evaluation on IDEA-Bench and DesignBench shows that APE achieves stronger alignment with improved efficiency. A user study with 128 participants on user-defined tasks demonstrates 19.8% higher perceived alignment without increased workload. Our work contributes a principled approach to prompting that offers an effective and efficient complement to the prevailing prompt-based interaction paradigm with text-to-image models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。