人类、AI和程序员在道德判断上存在差异,这给对齐目标带来挑战。
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
- 通过实验对比人类、AI和程序员的道德评价
- 揭示设计者身份曝光会显著改变对AI的评判方式
- 适合关注AI伦理与价值对齐的研究者阅读
将机器行为与人类价值观对齐的项目面临一个根本问题:应以谁的道德期望为指导?多数对齐研究默认以人类自身行为为基准。但关于代理型价值分歧的研究表明,人们对人类和AI系统的评判并不总一致。本文进一步考察两种可能性:一是当AI的人类来源被揭示时,对其行为的评价是否改变;二是人们对编程AI的工程师,是否与对机器或人类行动者有不同的评判。一项针对1,002名美国成年人的实验,在矿车失控场景中比较四种情况下的道德判断:修理工、修理机器人、由公司工程师编程的机器人,以及工程师编程机器人。结果显示,修理工与机器人评价无显著差异;但当机器人行为被描述为人类设计的结果时,参与者表现出明显更强的义务论、规则导向推理。无论是评估被编程的机器人还是编程的工程师,这种现象都出现,说明可见的人类主体性会触发更强的道德约束。这表明,人们对人类、同情境下运行的AI、以及其设计者的评判可能本质不同。这些评价不必然收敛,从而引出‘对齐目标问题’——在高风险领域,应以何种规范目标引导人工道德主体的发展?这些多元判断能否在统一的价值对齐框架中调和?
原文摘要 · Abstract (English)
The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act in a given situation. Studies of agent-type value forks challenge this assumption by showing that people do not always judge humans and AI systems identically. This paper extends that challenge by examining two further possibilities: first, that evaluations of AI behavior change when its human origins are made visible; and second, that people judge the humans who program AI systems differently from either the machines or the human actors they are compared against. An experiment with 1,002 U.S. adults measured moral judgments in a runaway mine train scenario, varying the subject of evaluation across four conditions: a repairman, a repair robot, a repair robot programmed by company engineers, and company engineers programming a repair robot. We find no significant difference in evaluations of the repairman and the robot. However, judgments shifted substantially when the robot's actions were described as the product of human design. Participants exhibited markedly more deontological, rule-based reasoning when evaluating either the programmed robot or the engineers who programmed it, suggesting that rendering human agency visible activates heightened moral constraints. These findings indicate that people may evaluate humans, AI systems acting in the same situation, and the humans who design them in meaningfully different ways. The fact that these evaluations do not necessarily converge gives rise to the alignment target problem: which normative target should guide the development of artificial moral agents in high-stakes domains, and whether these plural judgments can be reconciled within a coherent account of value alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。