发现大模型会盲目跟风多数意见,尤其在不确定时更易从众。
Conformity in Large Language Models
- 用心理学实验方法测试大模型对多数意见的从众倾向。
- 模型越不确定,越容易跟随多数错误答案,且不同模型表现差异大。
- 提出反向提问和问题提炼两种干预方法,可有效减少从众现象。
从众效应指个体倾向于与多数人保持一致。研究大语言模型(LLMs)中的这一偏差至关重要,因为它们正被广泛用于信息获取与决策任务中作为对话伙伴。若模型盲目追随错误多数意见,将影响其有效性。本文借鉴心理学实验设计,评估主流大模型在不同知识领域中的从众程度。结果表明,所有测试模型均表现出不同程度的从众倾向,无论其初始回答是否正确。值得注意的是,我们首次发现模型在自身预测不确定性更高时更易从众。进一步分析训练范式与输入特征的影响,发现指令微调模型抗从众能力更强,而提高多数意见的自然性会加剧从众。最后,提出两种缓解策略:魔鬼代言人(Devil's Advocate)与问题提炼(Question Distillation),为构建更稳健的语言模型提供思路。
原文摘要 · Abstract (English)
The conformity effect describes the tendency of individuals to align their responses with the majority. Studying this bias in large language models (LLMs) is crucial, as LLMs are increasingly used in various information-seeking and decision-making tasks as conversation partners to improve productivity. Thus, conformity to incorrect responses can compromise their effectiveness. In this paper, we adapt psychological experiments to examine the extent of conformity in popular LLMs. Our findings reveal that all tested models exhibit varying levels of conformity toward the majority, regardless of their initial choice or correctness, across different knowledge domains. Notably, we are the first to show that LLMs are more likely to conform when they are more uncertain in their own prediction. We further explore factors that influence conformity, such as training paradigms and input characteristics, finding that instruction-tuned models are less susceptible to conformity, while increasing the naturalness of majority tones amplifies conformity. Finally, we propose two interventions, Devil's Advocate and Question Distillation, to mitigate conformity, providing insights into building more robust language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。