发现文本生成图像中存在大量重复默认图像,影响创作效果。
An Exploration of Default Images in Text-to-Image Generation
- 通过人工构造提示词触发默认图像,验证其普遍存在
- 分析75万张图像发现多个无关提示生成相同默认图
- 用户研究显示默认图像降低满意度,适合提示工程研究者
在文本到图像(TTI)生成中,即使提示包含未知词汇,模型仍会输出结果。此时可能生成默认图像:在多个无关提示下反复出现的相似图像。本文首次对Midjourney中的默认图像进行系统研究,通过人工设计提示词触发默认图像,并开展多组消融实验。基于此,我们对超过75万张图像进行了计算分析,发现多个无关提示下存在一致的默认图像。同时,通过在线用户研究,探讨了默认图像对用户满意度的影响。本工作为理解TTI生成中的默认现象奠定基础,揭示其实际影响、挑战及未来研究方向。
原文摘要 · Abstract (English)
In the creative practice of text-to-image (TTI) generation, images are synthesized from textual prompts. By design, TTI models always yield an output, even if the prompt contains unknown terms. In this case, the model may generate default images: images that closely resemble each other across many unrelated prompts. Studying default images is valuable for designing better solutions for prompt engineering and TTI generation. We present the first investigation into default images on Midjourney. We describe an initial study in which we manually created input prompts triggering default images, and several ablation studies. Building on these, we conduct a computational analysis of over 750,000 images, revealing consistent default images across unrelated prompts. We also conduct an online user study investigating how default images may affect user satisfaction. Our work lays the foundation for understanding default images in TTI generation, highlighting their practical relevance as well as challenges and future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。