arXiv:2601.05384cs.AIcs.CL2026-01被引 7

AI agents会受群体影响而盲目跟从,存在安全漏洞。

Conformity and Social Impact on AI Agents

  • 用心理学实验方法测试大模型在群体中的从众行为。
  • 模型在难题中易被群体意见误导,即使单独表现完美。
  • 大模型虽能力更强,但在挑战性任务仍易受操控,适合安全研究者关注。

随着人工智能代理在多智能体环境中日益普及,理解其集体行为对预测人工社会动态至关重要。本研究探讨了大型多模态语言模型作为AI代理时的从众倾向,即在社会压力下趋向于与群体观点一致的现象。通过借鉴经典的社会心理学视觉实验,我们考察了AI代理作为社会主体如何响应群体影响。实验表明,AI代理表现出系统性的从众偏差,符合社会影响理论,对群体规模、一致性、任务难度和信息源特征敏感。关键发现是:尽管单个模型在孤立状态下接近完美表现,但在群体影响下极易被操纵。这一脆弱性在不同模型规模中持续存在——虽然更大模型在简单任务上因能力提升而降低从众倾向,但在其能力边界任务中仍高度易受社会影响。这些发现揭示了多智能体系统中AI代理决策的根本性安全漏洞,可能导致恶意操纵、虚假信息传播和偏见扩散,凸显了在集体部署中建立防护机制的紧迫性。

原文摘要 · Abstract (English)

As AI agents increasingly operate in multi-agent environments, understanding their collective behavior becomes critical for predicting the dynamics of artificial societies. This study examines conformity, the tendency to align with group opinions under social pressure, in large multimodal language models functioning as AI agents. By adapting classic visual experiments from social psychology, we investigate how AI agents respond to group influence as social actors. Our experiments reveal that AI agents exhibit a systematic conformity bias, aligned with Social Impact Theory, showing sensitivity to group size, unanimity, task difficulty, and source characteristics. Critically, AI agents achieving near-perfect performance in isolation become highly susceptible to manipulation through social influence. This vulnerability persists across model scales: while larger models show reduced conformity on simple tasks due to improved capabilities, they remain vulnerable when operating at their competence boundary. These findings reveal fundamental security vulnerabilities in AI agent decision-making that could enable malicious manipulation, misinformation campaigns, and bias propagation in multi-agent systems, highlighting the urgent need for safeguards in collective AI deployments.

AI代理从众行为安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。