arXiv:2504.14868cs.CV2025-04被引 20

通过对话逐步优化图像生成,让模糊提示也能产出精准结果。

Twin Co-Adaptive Dialogue for Progressive Image Generation

  • 用动态对话交互逐步修正图像,而非一次生成
  • 减少用户试错次数,提升图像质量与意图契合度
  • 适合需要精细调控的创意设计场景

现代文本到图像生成系统虽能产出高度逼真的视觉内容,但在处理用户提示中的固有歧义时仍表现不佳。本文提出Twin-Co框架,通过同步协同对话实现渐进式图像生成。系统首先根据用户提示生成基础图像,随后通过一系列同步对话交互,持续根据用户反馈调整优化图像。该协同适应机制能逐步消除歧义,更准确匹配用户意图。实验表明,Twin-Co不仅显著降低用户试错迭代次数,提升生成图像质量,还简化了各类应用场景下的创作流程。

原文摘要 · Abstract (English)

Modern text-to-image generation systems have enabled the creation of remarkably realistic and high-quality visuals, yet they often falter when handling the inherent ambiguities in user prompts. In this work, we present Twin-Co, a framework that leverages synchronized, co-adaptive dialogue to progressively refine image generation. Instead of a static generation process, Twin-Co employs a dynamic, iterative workflow where an intelligent dialogue agent continuously interacts with the user. Initially, a base image is generated from the user's prompt. Then, through a series of synchronized dialogue exchanges, the system adapts and optimizes the image according to evolving user feedback. The co-adaptive process allows the system to progressively narrow down ambiguities and better align with user intent. Experiments demonstrate that Twin-Co not only enhances user experience by reducing trial-and-error iterations but also improves the quality of the generated images, streamlining creative process across various applications.

图像生成对话系统人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。