arXiv:2605.25256cs.AI2026-05中稿 · ICML

比较不同组织中大模型决策过程对齐情况,发现对齐效果因组织和任务而异。

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

论文配图:Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts
图 1 · 摘自论文原文
  • 用决策政策捕捉法评估模型是否忠实复现组织决策逻辑而非仅结果一致
  • 在欧洲人权法院案件中,对齐度与准确率相关(r=0.85,p<0.001)
  • 在信贷决策中,模型抗拒采纳组织对保护属性的权重,高对齐未必理想

可调控的多元主义要求模型能忠实呈现特定视角。组织是这一需求的天然场景,因其需部署大模型做出反映自身政策的决策。然而现有研究多将视角固定于个人或人口群体。本文采用决策政策捕捉方法,评估大模型在组织环境中的过程对齐程度,即是否忠实复现组织的决策政策,而不仅是达成相同结论。研究发现对齐存在双重异质性:不同模型间基线对齐差异显著,且不随定价或通用基准性能变化;不同组织间对齐结构亦不同。在欧洲人权法院第6条判决中,过程对齐预测输出准确率(r=0.85,p<0.001),显式化组织历史决策政策可提升表现差的模型。在消费信贷决策中,整体对齐度低但变异大于输出准确率,模型抗拒采纳组织对受保护属性的权重。由于历史信贷决策可能包含歧视性模式,此处高对齐并非总是理想。因此过程级测量必不可少,同一程序可依据目标政策是否合乎伦理,用于校准或审计模型。决定对齐目标及对齐可行性与合理性,使组织对齐本身成为一个多元问题。

原文摘要 · Abstract (English)

Steerable pluralism requires a model to faithfully represent one specified perspective. Organizations are a natural setting for this demand, since they deploy LLMs to make decisions that must reflect their own policy. Yet, most existing work fixes that perspective at the level of individuals or demographic groups. We rely on a decision-policy capturing method to measure process alignment in organizational settings, assessing whether an LLM faithfully reproduces the organization's decision policy rather than merely reaching the same conclusions. We find heterogeneity along two axes. Across models, baseline alignment varies strongly and tracks neither pricing nor general benchmark performance. Across organizations, the structure of alignment changes. In ECHR Article 6 decisions, process alignment predicts output accuracy ($r = 0.85$, $p < .001$), and making the organization's past decision policy explicit improves poorly aligned models. In consumer credit decisions, process alignment is low overall but varies more than output accuracy, and the models resist adopting the organization's weighting of protected attributes. Because historical credit decisions encode potentially discriminatory patterns, higher alignment there is not always desirable. Process-level measurement is therefore necessary, and depending on whether the target policy is normatively desirable, the same procedure can calibrate or audit a model. Deciding which policy to align to, and whether higher alignment is feasible or desirable, makes organizational alignment a pluralistic problem in its own right.

大模型对齐组织决策过程对齐公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。