用户提示词越相似,生成图像越单调,需鼓励多样化表达
Exploring Language Patterns of Prompts in Text-to-Image Generation and Their Impact on Visual Diversity
- 分析六百万条提示词,发现用户随时间趋同于使用热门标签
- 重复提示占40%-50%,导致图像视觉多样性下降
- 提示词语言重复会加剧生成内容的同质化,影响公平性
在文本到图像(TTI)模型热度消退后,研究开始关注其社会技术动态。本研究分析了CivitAI平台上六百万条来自Civiverse数据集的提示词,覆盖七个月时间,将用户分为三类:持续重复者、偶尔重复者和非重复者。结果表明,随着用户参与度提升,提示词语言趋于同质化,大量采用社区流行标签与描述,重复提示占比达40%-50%。尽管语义相似性和主题偏好保持稳定,仍以常见主题和表面美学为主。通过Vendi分数量化视觉多样性,发现提示词词汇相似性与生成图像视觉相似性显著相关,语言重复强化了视觉单一性。该研究揭示用户行为对生成图像多样性的重要影响,超越模型固有偏见,呼吁开发促进语言与主题多样性的工具与实践,以推动更包容的AI生成内容。
原文摘要 · Abstract (English)
Following the initial excitement, Text-to-Image (TTI) models are now being examined more critically. While much of the discourse has focused on biases and stereotypes embedded in large-scale training datasets, the sociotechnical dynamics of user interactions with these models remain underexplored. This study examines the linguistic and semantic choices users make when crafting prompts and how these choices influence the diversity of generated outputs. Analyzing over six million prompts from the Civiverse dataset on the CivitAI platform across seven months, we categorize users into three groups based on their levels of linguistic experimentation: consistent repeaters, occasional repeaters, and non-repeaters. Our findings reveal that as user participation grows over time, prompt language becomes increasingly homogenized through the adoption of popular community tags and descriptors, with repeated prompts comprising 40-50% of submissions. At the same time, semantic similarity and topic preferences remain relatively stable, emphasizing common subjects and surface aesthetics. Using Vendi scores to quantify visual diversity, we demonstrate a clear correlation between lexical similarity in prompts and the visual similarity of generated images, showing that linguistic repetition reinforces less diverse representations. These findings highlight the significant role of user-driven factors in shaping AI-generated imagery, beyond inherent model biases, and underscore the need for tools and practices that encourage greater linguistic and thematic experimentation within TTI systems to foster more inclusive and diverse AI-generated content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。