发现大模型生成内容趋同,揭示‘人工蜂群’现象
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
- 构建26000条真实开放问题数据集,覆盖17类创意任务
- 多模型生成高度相似,单模型内部也反复输出相同内容
- 人类偏好差异大,但模型评估系统难以准确捕捉这种个体化偏好
语言模型在生成多样化、类人创意内容方面表现不佳,引发长期思想趋同的担忧。现有评估方法受限于狭窄任务或单一模型重复采样。本文提出Infinity-Chat,一个包含26,000条真实世界开放性用户提问的大规模数据集,涵盖6大类、17个子类的开放式问题,无唯一正确答案。基于此,首次系统研究了大模型在开放生成中的模式崩溃现象,揭示显著的“人工蜂群”效应:(1)单模型内部重复输出,(2)不同模型间生成结果高度同质。该数据集还包含31,250条人类标注,每例有25位独立评价者进行绝对评分与成对偏好判断。研究发现,尽管大模型、奖励模型及判别器整体质量相当,但在面对引发个体差异偏好的生成结果时,其评估与人类评分校准不足。总体而言,Infinity-Chat是首个系统研究真实开放问题下大模型生成行为的资源,为缓解人工智能长期安全风险提供关键洞见。
原文摘要 · Abstract (English)
Language models (LMs) often struggle to generate diverse, human-like creative content, raising concerns about the long-term homogenization of human thought through repeated exposure to similar outputs. Yet scalable methods for evaluating LM output diversity remain limited, especially beyond narrow tasks such as random number or name generation, or beyond repeated sampling from a single model. We introduce Infinity-Chat, a large-scale dataset of 26K diverse, real-world, open-ended user queries that admit a wide range of plausible answers with no single ground truth. We introduce the first comprehensive taxonomy for characterizing the full spectrum of open-ended prompts posed to LMs, comprising 6 top-level categories (e.g., brainstorm & ideation) that further breaks down to 17 subcategories. Using Infinity-Chat, we present a large-scale study of mode collapse in LMs, revealing a pronounced Artificial Hivemind effect in open-ended generation of LMs, characterized by (1) intra-model repetition, where a single model consistently generates similar responses, and more so (2) inter-model homogeneity, where different models produce strikingly similar outputs. Infinity-Chat also includes 31,250 human annotations, across absolute ratings and pairwise preferences, with 25 independent human annotations per example. This enables studying collective and individual-specific human preferences in response to open-ended queries. Our findings show that LMs, reward models, and LM judges are less well calibrated to human ratings on model generations that elicit differing idiosyncratic annotator preferences, despite maintaining comparable overall quality. Overall, INFINITY-CHAT presents the first large-scale resource for systematically studying real-world open-ended queries to LMs, revealing critical insights to guide future research for mitigating long-term AI safety risks posed by the Artificial Hivemind.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。