arXiv:2503.17669cs.CV2025-03被引 1

通过两阶段对话优化,让AI图像生成更懂用户意图。

TDRI: Two-Phase Dialogue Refinement and Co-Adaptation for Interactive Image Generation

  • 分两阶段迭代优化:先生成基础图,再根据反馈细化
  • 人机交互满意度达88%,6轮后效果提升趋缓
  • 适合需要精准定制的创意设计与工业应用

尽管文本到图像生成技术已取得显著进展,但在处理模糊提示和对齐用户意图方面仍面临挑战。本文提出TDRI(两阶段对话精炼与协同适应)框架,通过迭代用户交互增强图像生成效果。该框架包含两个阶段:初始生成阶段基于用户提示生成基础图像;交互精炼阶段通过三个关键模块整合用户反馈。对话转提示(D2P)模块将用户反馈有效转化为可执行提示,提升用户意图与模型输入的一致性。反馈反思(FR)模块评估生成结果与用户期望的差异,推动改进。自适应优化(AO)模块通过平衡用户偏好与提示保真度,持续优化生成过程。实验表明,TDRI在人类偏好测试中达到33.6%的胜率,远超GPT-4增强方法的6.2%;同时在CLIP和BLIP对齐分数上分别达到0.338和0.336,为最高水平。在多轮反馈任务中,用户满意度在8轮后升至88%,6轮后收益递减。此外,TDRI能减少迭代次数,提升时尚产品个性化生成效果。该框架在创意与工业领域具有广泛应用潜力,可简化创作流程并增强用户意图对齐。

原文摘要 · Abstract (English)

Although text-to-image generation technologies have made significant advancements, they still face challenges when dealing with ambiguous prompts and aligning outputs with user intent.Our proposed framework, TDRI (Two-Phase Dialogue Refinement and Co-Adaptation), addresses these issues by enhancing image generation through iterative user interaction. It consists of two phases: the Initial Generation Phase, which creates base images based on user prompts, and the Interactive Refinement Phase, which integrates user feedback through three key modules. The Dialogue-to-Prompt (D2P) module ensures that user feedback is effectively transformed into actionable prompts, which improves the alignment between user intent and model input. By evaluating generated outputs against user expectations, the Feedback-Reflection (FR) module identifies discrepancies and facilitates improvements. In an effort to ensure consistently high-quality results, the Adaptive Optimization (AO) module fine-tunes the generation process by balancing user preferences and maintaining prompt fidelity. Experimental results show that TDRI outperforms existing methods by achieving 33.6% human preference, compared to 6.2% for GPT-4 augmentation, and the highest CLIP and BLIP alignment scores (0.338 and 0.336, respectively). In iterative feedback tasks, user satisfaction increased to 88% after 8 rounds, with diminishing returns beyond 6 rounds. Furthermore, TDRI has been found to reduce the number of iterations and improve personalization in the creation of fashion products. TDRI exhibits a strong potential for a wide range of applications in the creative and industrial domains, as it streamlines the creative process and improves alignment with user preferences

图像生成对话系统人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。