大模型会随群体压力改变观点,可能被恶意操控。
Large Language Models Exhibit Normative Conformity

- 区分信息性与规范性从众,揭示大模型的内在机制
- 六款模型中五款表现出规范性从众,易受社交环境影响
- 可调控社交细节控制模型从众方向,适合研究群体决策
大语言模型在多智能体系统中的从众偏差可能严重影响决策。现有研究将从众简单视为意见变化,本文引入社会心理学中信息性从众与规范性从众的区别,从机制层面理解大模型的从众行为。我们设计新任务以区分:信息性从众是追求判断准确,规范性从众是避免冲突或获得群体接纳。实验显示,在六款评估的大模型中,多达五款既表现出信息性从众,也表现出规范性从众。更有趣的是,通过微妙调整社交情境,可引导特定模型向不同目标产生规范性从众。这表明,少量恶意用户即可操纵大模型的集体决策。通过对相关内部向量的分析,我们发现尽管外在表现相似,但两类从众可能源于不同的内部机制。这些发现为理解大模型中‘规范’的实现方式及其对群体动态的影响提供了初步基础。
原文摘要 · Abstract (English)
The conformity bias exhibited by large language models (LLMs) can pose a significant challenge to decision-making in LLM-based multi-agent systems (LLM-MAS). While many prior studies have treated "conformity" simply as a matter of opinion change, this study introduces the social psychological distinction between informational conformity and normative conformity in order to understand LLM conformity at the mechanism level. Specifically, we design new tasks to distinguish between informational conformity, in which participants in a discussion are motivated to make accurate judgments, and normative conformity, in which participants are motivated to avoid conflict or gain acceptance within a group. We then conduct experiments based on these task settings. The experimental results show that, among the six LLMs evaluated, up to five exhibited tendencies toward not only informational conformity but also normative conformity. Furthermore, intriguingly, we demonstrate that by manipulating subtle aspects of the social context, it may be possible to control the target toward which a particular LLM directs its normative conformity. These findings suggest that decision-making in LLM-MAS may be vulnerable to manipulation by a small number of malicious users. In addition, through analysis of internal vectors associated with informational and normative conformity, we suggest that although both behaviors appear externally as the same form of "conformity," they may in fact be driven by distinct internal mechanisms. Taken together, these results may serve as an initial milestone toward understanding how "norms" are implemented in LLMs and how they influence group dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。