arXiv:2511.10032cs.HCcs.AI2025-11被引 5

AI对齐需应对人类道德偏好随时间波动的问题,否则可能误导系统决策。

Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback

  • 通过多次问卷调查分析400多人对肾移植患者的道德判断变化
  • 约6%-20%的判断随时间改变,模型稳定性也显著下降
  • 忽视时间变化会导致AI预测能力下降,影响高风险领域应用

道德对齐方法旨在捕捉人类利益相关者的道德偏好并融入AI系统,但该过程假设道德偏好是静态的,而实际上它们常随时间演变。理想的对齐应区分合法的道德推理变化与由注意力缺陷或认知偏差等任意因素引起的波动。然而,当前主流对齐方法大多忽略偏好随时间的变化,这在医疗等高风险场景中尤为危险,可能导致系统失信并引发个体与社会性危害。本研究以肾移植分配为场景,从超过400名参与者在3-5轮次中对成对假想患者进行选择,发现平均而言,同一情景在不同时间点被不同回答的比例达6%-20%(表现为“响应不稳定性”);部分参与者的决策模型也出现显著时变(“模型不稳定性”)。简单AI模型的预测性能随响应与模型不稳定性的增加而下降,且性能随时间推移持续减弱。这些结果揭示了对齐对象本身的复杂性:当用户偏好显著变动时,必须重新思考‘对齐什么’这一根本问题。

原文摘要 · Abstract (English)

Alignment methods in moral domains seek to elicit moral preferences of human stakeholders and incorporate them into AI. This presupposes moral preferences as static targets, but such preferences often evolve over time. Proper alignment of AI to dynamic human preferences should ideally account for "legitimate" changes to moral reasoning, while ignoring changes related to attention deficits, cognitive biases, or other arbitrary factors. However, common AI alignment approaches largely neglect temporal changes in preferences, posing serious challenges to proper alignment, especially in high-stakes applications of AI, e.g., in healthcare domains, where misalignment can jeopardize the trustworthiness of the system and yield serious individual and societal harms. This work investigates the extent to which people's moral preferences change over time, and the impact of such changes on AI alignment. Our study is grounded in the kidney allocation domain, where we elicit responses to pairwise comparisons of hypothetical kidney transplant patients from over 400 participants across 3-5 sessions. We find that, on average, participants change their response to the same scenario presented at different times around 6-20% of the time (exhibiting "response instability"). Additionally, we observe significant shifts in several participants' retrofitted decision-making models over time (capturing "model instability"). The predictive performance of simple AI models decreases as a function of both response and model instability. Moreover, predictive performance diminishes over time, highlighting the importance of accounting for temporal changes in preferences during training. These findings raise fundamental normative and technical challenges relevant to AI alignment, highlighting the need to better understand the object of alignment (what to align to) when user preferences change significantly over time.

AI对齐道德偏好时间动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。