arXiv:2412.10419cs.CVcs.AI2024-12ICML被引 6

让AI理解用户逐步修改需求,动态优化图像生成。

Preference Adaptive and Sequential Text-to-Image Generation

  • 用强化学习根据用户反馈迭代优化提示词
  • 通过人类评分构建序列偏好数据集,提升交互体验
  • 适合需要多轮协作的创意设计场景

针对交互式文本到图像生成问题,我们设计了一个强化学习代理,通过一系列提示词扩展迭代改进图像以满足用户需求。利用人工评分者构建了新的序列偏好数据集,并结合大规模开源非序列数据集进行训练。采用期望最大化(EM)策略构建用户偏好与选择模型,识别出不同的用户偏好类型。随后,利用大型多模态语言模型(LMM)和基于价值的强化学习方法,为用户提供自适应且多样化的提示词扩展建议。所提出的偏好自适应与序列化文生图代理(PASTA)赋予文生图模型多轮交互能力,促进人机协同创作,缓解用户意图模糊或不完整的问题。通过人类评估验证,PASTA显著优于基线方法。同时,我们开源了序列评分者数据集及模拟用户-评分者交互环境,以支持未来面向用户的多轮文生图系统研究。

原文摘要 · Abstract (English)

We address the problem of interactive text-to-image (T2I) generation, designing a reinforcement learning (RL) agent which iteratively improves a set of generated images for a user through a sequence of prompt expansions. Using human raters, we create a novel dataset of sequential preferences, which we leverage, together with large-scale open-source (non-sequential) datasets. We construct user-preference and user-choice models using an EM strategy and identify varying user preference types. We then leverage a large multimodal language model (LMM) and a value-based RL approach to suggest an adaptive and diverse slate of prompt expansions to the user. Our Preference Adaptive and Sequential Text-to-image Agent (PASTA) extends T2I models with adaptive multi-turn capabilities, fostering collaborative co-creation and addressing uncertainty or underspecification in a user's intent. We evaluate PASTA using human raters, showing significant improvement compared to baseline methods. We also open-source our sequential rater dataset and simulated user-rater interactions to support future research in user-centric multi-turn T2I systems.

文生图强化学习交互设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。