arXiv:2504.13392cs.CVcs.HC2025-04被引 12

POET让AI图像生成更个性化,自动扩展创意维度,提升创作多样性。

POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation

  • 自动发现并扩展文本到图像模型的同质维度,扩大输出空间。
  • 用户反馈驱动个性化扩展,28人测试显示减少提示次数,提升满意度。
  • 适合创意设计、艺术生成等需要多样化灵感的场景。

当前先进的视觉生成AI工具在创意初期阶段具有巨大潜力,能生成高质量、前所未有的图像,并满足多样化的用户需求。然而,许多大规模文本到图像系统为通用性设计,导致输出趋于常规,限制了创意探索,且交互方式对新手不友好。由于创意用户常以多变、非预测的方式工作,亟需更多样化和个性化支持。我们提出POET,一个实时交互工具,能够(1)自动发现文本到图像生成模型中的同质维度;(2)扩展这些维度以丰富生成图像的空间;(3)根据用户反馈学习并个性化扩展。在四个创意任务领域的28名用户评估中,POET展现出更高的感知多样性,帮助用户以更少的提示达到满意结果,促使他们在共创过程中更深入地思考和反思多种可能产出。面向视觉创意,POET首次展示了未来文本到图像生成工具的交互方式如何更好地支持多元价值与用户在构思阶段的需求。

原文摘要 · Abstract (English)

State-of-the-art visual generative AI tools hold immense potential to assist users in the early ideation stages of creative tasks -- offering the ability to generate (rather than search for) novel and unprecedented (instead of existing) images of considerable quality that also adhere to boundless combinations of user specifications. However, many large-scale text-to-image systems are designed for broad applicability, yielding conventional output that may limit creative exploration. They also employ interaction methods that may be difficult for beginners. Given that creative end users often operate in diverse, context-specific ways that are often unpredictable, more variation and personalization are necessary. We introduce POET, a real-time interactive tool that (1) automatically discovers dimensions of homogeneity in text-to-image generative models, (2) expands these dimensions to diversify the output space of generated images, and (3) learns from user feedback to personalize expansions. An evaluation with 28 users spanning four creative task domains demonstrated POET's ability to generate results with higher perceived diversity and help users reach satisfaction in fewer prompts during creative tasks, thereby prompting them to deliberate and reflect more on a wider range of possible produced results during the co-creative process. Focusing on visual creativity, POET offers a first glimpse of how interaction techniques of future text-to-image generation tools may support and align with more pluralistic values and the needs of end users during the ideation stages of their work.

文本生成图像生成个性化创意辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。