从噪声偏好中学习多目标奖励,实现更符合人类复杂价值观的对齐。
Learning Pareto-Optimal Rewards from Noisy Preferences: A Framework for Multi-Objective Inverse Reinforcement Learning
- 将人类偏好建模为隐含的向量奖励函数,构建多目标逆强化学习框架。
- 在噪声偏好下可恢复ε近似帕累托前沿,样本复杂度有理论保证。
- 算法收敛且适用于高维多价值场景,适合对齐复杂生成智能体行为。
随着生成式智能体能力提升,如何使其行为与复杂的隐含人类价值观对齐仍是根本挑战。现有方法常将人类意图简化为标量奖励,忽略了人类反馈的多维度特性。本文提出一种基于偏好的多目标逆强化学习(MO-IRL)理论框架,将人类偏好建模为潜在的向量值奖励函数。我们形式化了从噪声偏好查询中恢复帕累托最优奖励表示的问题,并建立了识别底层多目标结构的条件。推导出恢复ε-近似帕累托前沿的紧致样本复杂度界,并引入后悔度量来量化该多目标设定下的次优性。此外,我们提出一个可证明收敛的算法,用于基于偏好推断的奖励锥进行策略优化。结果弥合了实际对齐技术与理论保证之间的差距,为在高维、多元价值环境中学习对齐行为提供了原则性基础。
原文摘要 · Abstract (English)
As generative agents become increasingly capable, alignment of their behavior with complex human values remains a fundamental challenge. Existing approaches often simplify human intent through reduction to a scalar reward, overlooking the multi-faceted nature of human feedback. In this work, we introduce a theoretical framework for preference-based Multi-Objective Inverse Reinforcement Learning (MO-IRL), where human preferences are modeled as latent vector-valued reward functions. We formalize the problem of recovering a Pareto-optimal reward representation from noisy preference queries and establish conditions for identifying the underlying multi-objective structure. We derive tight sample complexity bounds for recovering $ε$-approximations of the Pareto front and introduce a regret formulation to quantify suboptimality in this multi-objective setting. Furthermore, we propose a provably convergent algorithm for policy optimization using preference-inferred reward cones. Our results bridge the gap between practical alignment techniques and theoretical guarantees, providing a principled foundation for learning aligned behaviors in a high-dimension and value-pluralistic environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。