开发者对道德AI投票设计的三重选择,直接影响最终决策公平性。
Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers

- 通过特征选择、投票者筛选和问题表述三步,开发者暗中塑造道德偏好。
- 不同情境下道德特征差异大,政治立场影响约1/3特征的倾向,方向可反转。
- 问题措辞能扩大或缩小意识形态差距,最高达1个量表点,影响判断基础。
随着人工智能在社会中做出越来越多具有道德意味的决策,一种应对方式是进行道德偏好收集。该方法通过向参与者提出假设性困境并汇总投票结果,训练出可在大规模应用的政策模型。但在任何投票前,开发者已在道德AI偏好收集流程中做出三个关键选择:特征范围界定、投票者抽样与问题表述方式。这些选择常不透明、无记录,被当作技术细节而非价值判断。本文在三项实际部署场景(即AI肾脏分配、模拟缺席员工的AI代理、生成式AI对逝者形象再现)中,考察了这一流程的三个阶段。第一,道德相关特征随场景变化,表明特征框架不可跨领域通用。第二,约三分之一特征的偏好受政治意识形态影响,部分差异甚至方向逆转,投票群体的政治构成会显著改变整体偏好分布。第三,问题表述方式可使意识形态差距扩大或缩小最多达一个量表点,同时改变道德基础与判断之间的关联模式。研究结论表明,仅靠投票聚合无法实现公平或透明的AI对齐;至少每个环节都需审计并公开披露。
原文摘要 · Abstract (English)
As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. In this approach, researchers poll participants on hypothetical dilemmas and use the aggregated votes to train a policy that an AI model then applies at scale. Before any vote is cast, developers make three key choices in the moral AI elicitation pipeline: feature scoping, voter sampling, and question framing. In other words, they decide which features go to a vote, which voters to include, and how to present the question. These choices are often opaque, undocumented, and treated as technical details rather than normative ones. We examine each of these choices within a common empirical study and show that each can shape the preferences produced by moral AI elicitation. Across two phases (N = 809) in three deployment contexts (i.e., AI kidney allocation, AI agents simulating absent workers, and generative AI depictions of the deceased), we examine the three main stages of the moral AI elicitation pipeline. First, morally relevant features shift across contexts. This suggests that feature schemas should not be assumed to transfer across deployment domains. Second, preferences differ by political ideology for roughly one-third of features, with some differences reversing direction. The ideological composition of the voter pool can therefore affect the resulting aggregated preference profile. Third, the wording of the elicitation question can narrow or widen ideological gaps by up to a full scale point. The framing conditions also change how moral foundations are associated with participants' judgments. Taken together, these findings suggest that voting-based alignment cannot deliver fair or transparent AI by aggregation alone; at minimum, each stage of the moral AI elicitation pipeline should be audited and disclosed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。