arXiv:2503.10215cs.AIcs.GT2025-03被引 2

提出自适应偏好聚合方法,解决AI对齐中人类偏好整合难题。

Adaptive Preference Aggregation

  • 基于瓮模型设计上下文感知的偏好聚合策略
  • 继承最大彩票解法特性,满足康多塞一致性
  • 适用于多维场景的AI对齐与推荐系统

AI对齐问题——确保人工智能系统行为符合人类价值观——已成为基础模型和推荐系统发展中的关键挑战。当前主流的基于人类反馈的强化学习(RLHF)在整合多样人类偏好方面存在已知理论局限。社会选择理论提供了偏好聚合的框架,但其未针对人工智能常见的多维应用场景设计。本文借鉴近期发表的瓮过程思想,提出一种能随用户上下文自适应调整的偏好聚合策略,该策略继承了最大彩票解法(maximal lottery)的优良性质,是一种满足康多塞一致性的解决方案。

原文摘要 · Abstract (English)

AI alignment, the challenge of ensuring AI systems act in accordance with human values, has emerged as a critical problem in the development of systems such as foundation models and recommender systems. Still, the current dominant approach, reinforcement learning with human feedback (RLHF) faces known theoretical limitations in aggregating diverse human preferences. Social choice theory provides a framework to aggregate preferences, but was not developed for the multidimensional applications typical of AI. Leveraging insights from a recently published urn process, this work introduces a preference aggregation strategy that adapts to the user's context and that inherits the good properties of the maximal lottery, a Condorcet-consistent solution concept.

AI对齐偏好聚合社会选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。