arXiv:2601.01904cs.LGcs.AI2026-01被引 1

发现偏好学习中特征相关噪声会严重拖垮主流鲁棒方法

Evaluating Feature Dependent Noise in Preference-based Reinforcement Learning

  • 提出轨迹特征、相似度、边际、语言模型等四类特征依赖噪声
  • 在DMControl和Meta-World上验证,多数场景下无去噪方法反而更优
  • 揭示语言模型噪声与特征噪声特性相似,提示需更真实建模人类偏好

基于偏好的强化学习(PbRL)近年来受到关注,因其适用于奖励函数难以获取的复杂任务。然而,当偏好来自非完美教师时,常伴随不确定性与噪声。以往研究多关注均匀分布的噪声类型,且缺乏与观测特征的关联。本文首次形式化了目标特征依赖噪声,并提出轨迹特征噪声、轨迹相似性噪声、边际依赖噪声及语言模型噪声等变体。我们在复杂连续控制任务(来自DMControl和Meta-World)中评估了这类噪声的影响。实验表明,在某些特征依赖噪声设置下,当前最先进的抗噪PbRL方法性能显著下降;而未显式去噪的PbRL方法反而在多数场景中表现更优。此外,我们发现语言模型产生的噪声具有类似特征依赖性,可模拟真实人类偏好,提示应加强对特征依赖噪声的鲁棒学习研究。

原文摘要 · Abstract (English)

Learning from Preferences in Reinforcement Learning (PbRL) has gained attention recently, as it serves as a natural fit for complicated tasks where the reward function is not easily available. However, preferences often come with uncertainty and noise if they are not from perfect teachers. Much prior literature aimed to detect noise, but with limited types of noise and most being uniformly distributed with no connection to observations. In this work, we formalize the notion of targeted feature-dependent noise and propose several variants like trajectory feature noise, trajectory similarity noise, margin dependent noise, and Language Model noise. We evaluate feature-dependent noise, where noise is correlated with certain features in complex continuous control tasks from DMControl and Meta-world. Our experiments show that in some feature-dependent noise settings, the state-of-the-art noise-robust PbRL method's learning performance is significantly deteriorated, while PbRL method with no explicit denoising can surprisingly outperform noise-robust PbRL in the majority of settings. We also find language models' noise exhibits similar characteristics to feature-dependent noise, thereby simulating realistic humans and call for further study in learning with feature-dependent noise robustly.

强化学习偏好学习噪声鲁棒性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。