arXiv:2506.12529cs.LGcs.AI2025-06被引 2

用相似性对齐偏好,让模型更抗错误标注且适应多种反馈。

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

  • 通过对比学习构建优选样本的隐空间表示,用相似度计算奖励。
  • 在含噪声标签数据上表现更稳定,显著优于基线模型(p<0.01)。
  • 适用于非专家标注,适合实际场景中复杂反馈的强化学习任务。

基于偏好的强化学习(PbRL)旨在通过多样方法对齐模型与人类意图,以减轻奖励工程负担。然而,以往多数工作未考虑标注者错误这一现实问题,尤其当标注者为非专家或时间紧张时。本文提出相似性作为奖励对齐(SARA),一种简单且鲁棒的对比框架,既能抵抗噪声标签,又能适应多样的反馈形式。SARA通过学习优选样本的隐表示,并将奖励定义为样本与该表示的相似度。在具有不同真实噪声率的偏好数据上,我们展示了其在连续控制离线强化学习基准上的竞争性且更稳定的性能,相比基线有统计显著提升(Wilcoxon符号秩检验,p < 0.01)。此外,我们以环境奖励相关性作为偏好对齐程度的代理指标,结果表明,无论噪声率如何,SARA计算的奖励均表现出更高相关性。

原文摘要 · Abstract (English)

Preference-based Reinforcement Learning (PbRL) entails a variety of approaches for aligning models with human intent to alleviate the burden of reward engineering. However, most previous PbRL work has not investigated the robustness to labeler errors, inevitable with labelers who are non-experts or operate under time constraints. We introduce Similarity as Reward Alignment (SARA), a simple contrastive framework that is both resilient to noisy labels and adaptable to diverse feedback formats. SARA learns a latent representation of preferred samples and computes rewards as similarities to the learned latent. On preference data with varying realistic noise rates, we demonstrate competitive and more stable performance on continuous control offline RL benchmarks, with statistically significant improvements over baselines (Wilcoxon signed-rank, p < 0.01). We also compute correlation to the environment rewards as a proxy for measuring alignment to the underlying preference criteria. We show that the SARA computed rewards display higher correlation across noise rates compared to baselines.

强化学习偏好对齐噪声鲁棒对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。