arXiv:2501.15167cs.CV2025-01被引 37

通过人机协同优化意图理解,减少图像生成中反复修改提示的次数。

Enhancing Intent Understanding for Ambiguous prompt: A Human-Machine Co-Adaption Strategy

  • 以用户提示与图像间的互信息为优化目标,实现人机协同调参。
  • 实验表明新方法显著降低用户需多次调整提示的频率。
  • 适用于非专业用户,提升图像生成系统的易用性与意图契合度。

当前图像生成系统虽能产出高质量图像,但在面对模糊用户提示时,难以准确理解真实意图,导致用户需多次修改提示才能满意。现有方法多聚焦于优化提示本身,但模型仍难把握用户真实需求,尤其对非专业用户而言。本研究旨在优化视觉参数调整过程,使系统更友好、更懂用户。提出一种人机协同适应策略,以用户提示与待修改图像之间的互信息作为优化目标,提升系统对用户需求的适应能力。实验结果表明,改进模型可显著减少多轮调整的需求。同时,我们构建了包含多轮对话的标注数据集,涵盖提示、图像及用户意图。相关数据集与标注工具将公开共享。

原文摘要 · Abstract (English)

Current image generation systems produce high-quality images but struggle with ambiguous user prompts, making interpretation of actual user intentions difficult. Many users must modify their prompts several times to ensure the generated images meet their expectations. While some methods focus on enhancing prompts to make the generated images fit user needs, the model is still hard to understand users' real needs, especially for non-expert users. In this research, we aim to enhance the visual parameter-tuning process, making the model user-friendly for individuals without specialized knowledge and better understand user needs. We propose a human-machine co-adaption strategy using mutual information between the user's prompts and the pictures under modification as the optimizing target to make the system better adapt to user needs. We find that an improved model can reduce the necessity for multiple rounds of adjustments. We also collect multi-round dialogue datasets with prompts and images pairs and user intent. Various experiments demonstrate the effectiveness of the proposed method in our proposed dataset. Our dataset and annotation tools will be available.

图像生成人机协同意图理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。