arXiv:2606.21550cs.AI2026-06中稿 · publication in ACM…

用社会选择理论解决人类反馈中的意见冲突问题

AI Alignment From Social Choice Perspectives

  • 从社会选择视角分析人类反馈聚合机制
  • 揭示反馈冲突导致的模型对齐失效模式
  • 为处理分歧提供系统化设计思路,适合对齐研究者

基于人类反馈的对齐方法利用人类对模型输出的判断来引导预训练语言模型的行为。当这些判断反映不同人对理想行为的分歧时,学习到的目标函数即为模型应偏好内容的综合判定。本文综述了近期从社会选择理论角度研究该聚合问题的工作。通过社会选择视角,我们阐明了反馈聚合层可能存在的失效模式,并揭示了以明确、严谨方式处理分歧的更广泛设计空间。

原文摘要 · Abstract (English)

Alignment from human feedback uses human judgments about model outputs to steer the behavior of language models after pretraining. When those judgments reflect conflicting views of desirable behavior, the learned objective becomes an aggregate determination of what the model should prefer. We survey recent work that has studied this aggregation problem through the lens of social choice theory. We illustrate how the social choice perspective helps identify failure modes in the feedback aggregation layer and reveals a broader design space for handling disagreement in explicit and principled ways.

对齐社会选择人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。