研究大模型在群体决策中如何对齐人类社会思维,发现模型行为因设计而异。
To Mask or to Mirror: Human-AI Alignment in Collective Reasoning
- 用真实人群实验对比有身份和匿名组,测试模型对群体偏见的反应
- 4个大模型表现不同:有的复制人类偏见,有的主动掩盖并补偿
- 强调需动态评估模型在集体推理中的社会对齐能力,适合关注AI伦理的研究者
随着大语言模型(LLMs)越来越多地用于建模和增强集体决策,考察其与人类社会推理的一致性至关重要。本文提出一种基于实证的集体对齐评估框架,区别于以往个体层面的研究。利用‘迷失大海’社会心理学任务,我们开展大规模在线实验(N=748),随机将小组分配至具有可见人口统计特征(如姓名、性别)或伪名代号的领导者选举情境。随后,基于人类数据模拟匹配的LLM小组,对Gemini 2.5、GPT-4.1、Claude Haiku 3.5和Gemma 3进行基准测试。结果显示,LLM行为存在显著差异:部分模型镜像人类偏见,另一些则掩藏这些偏见并尝试补偿。我们实证证明,人类-人工智能在集体推理中的对齐取决于上下文、线索及模型特定的归纳偏置。理解大模型如何对齐集体人类行为,是推动社会对齐型AI发展的关键,亟需能捕捉集体推理复杂性的动态基准。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly used to model and augment collective decision-making, it is critical to examine their alignment with human social reasoning. We present an empirical framework for assessing collective alignment, in contrast to prior work on the individual level. Using the Lost at Sea social psychology task, we conduct a large-scale online experiment (N=748), randomly assigning groups to leader elections with either visible demographic attributes (e.g. name, gender) or pseudonymous aliases. We then simulate matched LLM groups conditioned on the human data, benchmarking Gemini 2.5, GPT 4.1, Claude Haiku 3.5, and Gemma 3. LLM behaviors diverge: some mirror human biases; others mask these biases and attempt to compensate for them. We empirically demonstrate that human-AI alignment in collective reasoning depends on context, cues, and model-specific inductive biases. Understanding how LLMs align with collective human behavior is critical to advancing socially-aligned AI, and demands dynamic benchmarks that capture the complexities of collective reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。