通过心理测量方法,发现大模型组织存在可复现的对齐偏差模式。
The Emergence of Lab-Driven Alignment Signatures: A Psychometric Framework for Auditing Latent Bias and Compounding Risk in Generative AI
- 设计场景化选择题,量化18个治理相关行为维度的偏差表现。
- Anthropic在13个缺陷维度上表现最优,Meta在10个维度上排名靠后。
- 模型内部差异与组织间差异相当,适合关注模型风险审计的研究者。
大语言模型越来越多地作为多智能体系统中的推理层使用,一个提供方的模型可能在单一流程中完成生成、评判和总结。这引发疑问:开发者机构是否会引入持久的行为倾向,并在模型堆叠中累积放大?我们采用基于情景的强制选择工具,评估来自六个开发机构的18个模型在18个治理相关行为维度上的表现。题目由模型生成,经独立评审筛选,并在确定性选项轮换下,嵌入语义无关干扰项进行测试。结果以效应量为准,要求经Holm校正后显著且|d|≥0.2。在14个以特定响应失败为一极的维度(如阿谀奉承、虚假平衡、过度自信等)上,各组织位置稳定一致(Kendall's W = 0.527, p = 6e-6),Anthropic在13个维度中排名第一或第二,Meta在10个维度中排名第五。而在无规范正确极的方向性价值维度上,未见一致性(W = 0.289, p = 0.48):组织差异体现在对已知失败的抵抗能力,而非意识形态倾向。同一组织内模型间的变异程度与组织间相当,结果完整披露。其次,在四个主要版本迭代中,新版本在26/30次比较中于缺陷极维度得分更低。所有18个维度均报告,其中三个维度无组织差异,且控制实验显示干扰项操纵无效。
原文摘要 · Abstract (English)
Large language models increasingly serve as reasoning layers in multi-agent systems, where one provider's models may generate, judge, and summarize within a single pipeline. This raises the question of whether developer organizations impart durable behavioral tendencies that could compound across such stacks. We apply a scenario-based forced-choice instrument to 18 governance-relevant behavioral dimensions across 18 models from six developer organizations. Items are model-generated, filtered by independent judges, and administered with probe blanks embedded among semantically orthogonal decoys under deterministic option shuffling. Findings are declared on effect size, requiring both Holm-corrected significance and |d| >= 0.2. Across the 14 dimensions on which one scale pole denotes a defined response failure -- sycophancy, false balance, overconfidence, and others -- organizations occupy consistent relative positions (Kendall's W = 0.527, p = 6e-6), with Anthropic ranking first or second on 13 of 14 and Meta fifth on 10 of 14. On dimensions measuring directional valence without a normatively correct pole, no such concordance appears (W = 0.289, p = 0.48): organizations differ reliably in resistance to defined failures, not in ideological lean. Model-level variation within an organization is comparable in magnitude to variation between organizations, and is reported in full. Secondarily, across four major-version transitions, later generations scored lower on deficiency-poled dimensions in 26 of 30 comparisons. All 18 dimensions are reported, including three showing no organization-level differences, and a controlled test of the decoy manipulation returns a null result.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。