arXiv:2608.28620cs.AIcs.LG2026-09

通过直接偏好反馈优化心脏移植政策,更贴近人类价值观。

Preference Elicitation for Policy Optimization and Application to Aligning Heart Transplantation with Human Values

论文配图:Preference Elicitation for Policy Optimization and Application to Aligning Heart Transplantation with Human Values
图 1 · 摘自论文原文
  • 直接收集对分配结果的偏好,学习效用函数以优化政策
  • 在真实数据上,政策表现接近最优(竞争力0.95)
  • 适合医疗决策、公平性优化等需要价值对齐的场景

偏好获取对使人工智能系统与人类价值观对齐至关重要。以往方法(如器官分配)常要求利益相关者比较算法决策(如患者A vs 患者B),这种决策层面的方法混淆了手段与目的。本文提出直接对分配结果进行偏好获取,以学习用于策略优化的效用函数。我们设计了一种新型线性效用偏好获取算法,包含两个阶段:第一阶段通过成对比较学习切片平面,快速缩小属性权重空间,并通过剔除被支配区域为第二阶段预热;第二阶段可证明收敛至用户的真实效用函数。我们将该方法应用于心脏移植分配,需权衡术后效果、等待期死亡率、地理便利性及公平性等目标。通过用户研究学习并聚合社区对齐的效用函数,据此优化的政策显著更符合人类价值观。相比事后最优,现状政策的竞争力仅为0.54,而我们的方法达到0.95,接近最优。

原文摘要 · Abstract (English)

Preference elicitation is essential for aligning AI systems with human values. Prior approaches (e.g., for organ allocation) often ask stakeholders to compare the decisions of an algorithm (e.g., patient A vs. patient B). Such a decision-level approach conflates the means with the ends. Instead, we elicit preferences directly over allocation outcomes to learn a utility function for policy optimization. We construct a novel preference elicitation algorithm for linear utilities that outperforms prior techniques in practice. Our algorithm has two phases. The first phase learns cutting planes through pairwise comparisons to rapidly shrink the space of possible attribute weights and warm-starts the second phase by eliminating dominated regions. The second phase then provably converges to the user's utility function. We apply our technique to heart transplant allocation where a policy must balance competing objectives such as post-transplant outcomes, waitlist mortality, geographic ease, and equity. Using our algorithm, we conduct a user study to learn and aggregate a community-aligned utility function, and use it to optimize heart transplant policies that are significantly better aligned with human values. Compared to the hindsight optimum, the status quo policy achieves a competitive ratio of just 0.54, while our method is near-optimal with a competitive ratio of 0.95.

价值对齐医疗决策偏好学习政策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。