arXiv:2409.07464cs.CVcs.AI2024-09

让AI通过对话不断优化图像生成,更懂用户真实意图。

Reflective Human-Machine Co-adaptation for Enhanced Text-to-Image Generation Dialogue System

  • AI与用户多轮对话,动态调整生成策略。
  • 基于用户反馈优化决策,提升结果匹配度。
  • 适合希望精准生成图像的普通用户。

当前图像生成系统虽能产出高质量图像,但用户提示常含模糊性,导致系统难以准确理解其真实意图。为此,需通过多轮交互来澄清需求,但频繁交互带来不可预测的成本,限制了非专业用户的使用和模型潜力发挥。本研究提出一种反思式人机协同适应策略RHM-CAS:外部层面,智能体通过有意义的语言对话反思并优化生成图像;内部层面,基于用户偏好优化生成策略,使最终输出更贴近用户期望。在多种任务上的实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Today's image generation systems are capable of producing realistic and high-quality images. However, user prompts often contain ambiguities, making it difficult for these systems to interpret users' potential intentions. Consequently, machines need to interact with users multiple rounds to better understand users' intents. The unpredictable costs of using or learning image generation models through multiple feedback interactions hinder their widespread adoption and full performance potential, especially for non-expert users. In this research, we aim to enhance the user-friendliness of our image generation system. To achieve this, we propose a reflective human-machine co-adaptation strategy, named RHM-CAS. Externally, the Agent engages in meaningful language interactions with users to reflect on and refine the generated images. Internally, the Agent tries to optimize the policy based on user preferences, ensuring that the final outcomes closely align with user preferences. Various experiments on different tasks demonstrate the effectiveness of the proposed method.

图像生成人机交互对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。