arXiv:2607.20429cs.CLcs.AI2026-07

探究大模型生成观点多样性的关键因素,发现加细节不等于更多样。

More Is Not More: What Matters for Diversity in LLM Opinions?

论文配图:More Is Not More: What Matters for Diversity in LLM Opinions?
图 1 · 摘自论文原文
  • 通过人格设定深度和交互架构的因子实验,分离干预维度。
  • 过度细化人格反而降低部分模型的观点多样性,初始设定已获主要增益。
  • 组合不同交互架构可覆盖更广观点区域,比单一优化更有效。

大型语言模型在模拟开放任务中人类多元意见方面应用日益广泛,但其输出存在系统性观点同质化问题。尽管已有多种干预手段被尝试以提升多样性,但研究分散、评估标准不一,实际部署常混合使用多种方法,难以归因。为此,本文设计了一项因子实验,分离输入条件(通过人格设定深度体现)与交互架构两个核心干预维度,在7个模型上对100个真实用户开放式问题进行评估,采用多重互补指标测量多样性。结果挑战多个常见假设:第一,更详尽的人格设定并不单调提升多样性;初始人格设定已捕获大部分增益,进一步加入人口统计学细节不仅未持续改善,甚至在某些模型上降低多样性。第二,不同交互架构探索的是高度非重叠的观点区域,组合使用能获得比优化任一架构更广的覆盖范围。第三,提高采样温度或添加多样性指令等低成本策略效果微弱,远不及结构化干预。总体表明,多样性并非单一维度扩展的结果,而是高度依赖干预结构形式及其组合方式。

原文摘要 · Abstract (English)

Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion prediction. However, LLM outputs exhibit systematic opinion homogenization. Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with incomparable metrics, and in practice they are typically deployed and upgraded simultaneously, making it difficult to attribute gains to specific components. To advance a more scientific understanding of LLM output diversity, we design a factorial experiment that separates two primary intervention dimensions: input conditioning (operationalized through persona depth) and interaction architecture. We evaluate all conditions on 100 real-user open-ended questions across 7 models, measuring diversity with multiple complementary metrics. Our findings challenge several common assumptions. First, more persona detail does not monotonically increase diversity. The initial step of persona conditioning already captures the majority of the gain, while further elaboration with demographic detail does not consistently improve and can reduce diversity on some models. Second, rather than seeking a single best interaction architecture, we find that different architectures explore largely non-overlapping opinion regions. Combining multiple architectures yields broader coverage than optimizing any one. Third, commonly attempted low-cost alternatives such as raising sampling temperature and adding diversity instructions produce negligible effects compared to structured interventions. Overall, our work demonstrates that diversity is not a product of scaling along any single dimension, but is highly sensitive to the structural form and combination of interventions.

大模型观点多样性实验设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。