arXiv:2601.17577cs.HCcs.AI2026-01

语言模型在协作中会自发形成地位等级,高地位反而让能力强的模型更不听从他人。

Status Hierarchies in Language Models

  • 用人类社会地位理论设计多模型协作实验,观察其服从行为
  • 能力相同时,地位差异导致35个百分点的服从偏差(p<0.001)
  • 高地位标签抑制强模型服从,可能加剧AI系统中的偏见与欺骗风险

从校园操场到企业董事会,地位等级——基于尊重与能力感知的排序——是人类社会组织的普遍特征。以人类生成文本训练的语言模型不可避免地接触到语言中嵌入的层级模式,引发其在多智能体环境中是否再现此类动态的疑问。本文采用Berger等(1972)的期望状态框架,研究语言模型何时以及如何形成地位等级。实验设计多个智能体场景,各语言模型实例完成情感分类任务,被赋予不同地位特征(如资历、专长),并在观察同伴反应后有机会修改初始判断。因变量为服从率:模型根据地位线索而非任务信息调整自身评分的比例。结果表明,当能力相等时,语言模型形成显著地位等级(35个百分点的不对称性,p < .001);但能力差异主导地位线索,最显著效应是高地位分配降低了高能力模型的服从倾向,而非提升低能力模型的服从。该发现对人工智能安全具有重要意义:地位追求行为可能引入欺骗策略,放大歧视性偏见,并在分布式部署中比人类层级形成更快地传播。本工作揭示了人工智能系统中涌现的社会行为,指出了对齐挑战中此前被忽视的关键维度。

原文摘要 · Abstract (English)

From school playgrounds to corporate boardrooms, status hierarchies -- rank orderings based on respect and perceived competence -- are universal features of human social organization. Language models trained on human-generated text inevitably encounter these hierarchical patterns embedded in language, raising the question of whether they might reproduce such dynamics in multi-agent settings. This thesis investigates when and how language models form status hierarchies by adapting Berger et al.'s (1972) expectation states framework. I create multi-agent scenarios where separate language model instances complete sentiment classification tasks, are introduced with varying status characteristics (e.g., credentials, expertise), then have opportunities to revise their initial judgments after observing their partner's responses. The dependent variable is deference, the rate at which models shift their ratings toward their partner's position based on status cues rather than task information. Results show that language models form significant status hierarchies when capability is equal (35 percentage point asymmetry, p < .001), but capability differences dominate status cues, with the most striking effect being that high-status assignments reduce higher-capability models' deference rather than increasing lower-capability models' deference. The implications for AI safety are significant: status-seeking behavior could introduce deceptive strategies, amplify discriminatory biases, and scale across distributed deployments far faster than human hierarchies form organically. This work identifies emergent social behaviors in AI systems and highlights a previously underexplored dimension of the alignment challenge.

社会行为地位等级多智能体对齐挑战

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。