让机器人根据环境自动调整行为优先级,提升学习效率。
Learning Contextually-Adaptive Rewards via Calibrated Features
- 分离上下文敏感性与固定偏好,显式建模特征重要性变化
- 仅需5-10次交互即可达到基线10倍的样本效率
- 适合个性化智能体训练,用户可轻松表达情境偏好
从人类反馈中学习奖励函数的核心挑战在于:理想行为会随上下文动态变化。例如,炉子变热后,机器人应更关注远离炉灶。我们发现,高层偏好(如安全优于效率)通常恒定,但特征的显著性——即其重要程度——会随情境改变。例如,炉温升高会增强距离特征的权重,而非改变安全偏好本身。此外,这些上下文效应在不同任务间重复出现,亟需可迁移的表征来编码。现有多任务与元学习方法虽能同时学习表征和偏好,但仅隐式捕捉上下文影响,且需大量数据才能区分上下文效应与任务特异性偏好。本文提出显式分离上下文依赖的特征显著性与上下文无关的偏好,引入校准特征(calibrated features)——模块化表示,专门捕捉上下文对特征重要性的影响,并设计专用配对比较查询以高效解耦显著性与偏好。模拟实验表明,本方法在低数据场景下(5–10次查询)相比基线减少10倍偏好查询量,性能提升最高达15%。真人用户研究(N=12)验证了该方法能有效传达个人情境偏好,实现可适应、个性化的奖励学习。
原文摘要 · Abstract (English)
A key challenge in reward learning from human input is that desired agent behavior often changes based on context. For example, a robot must adapt to avoid a stove once it becomes hot. We observe that while high-level preferences (e.g., prioritizing safety over efficiency) often remain constant, context alters the $\textit{saliency}$--or importance--of reward features. For instance, stove heat changes the relevance of the robot's proximity, not the underlying preference for safety. Moreover, these contextual effects recur across tasks, motivating the need for transferable representations to encode them. Existing multi-task and meta-learning methods simultaneously learn representations and task preferences, at best $\textit{implicitly}$ capturing contextual effects and requiring substantial data to separate them from task-specific preferences. Instead, we propose $\textit{explicitly}$ modeling and learning context-dependent feature saliency separately from context-invariant preferences. We introduce $\textit{calibrated features}$--modular representations that capture contextual effects on feature saliency--and present specialized paired comparison queries that isolate saliency from preference for efficient learning. Simulated experiments show our method improves sample efficiency, requiring 10x fewer preference queries than baselines to achieve equivalent reward accuracy, with up to 15% better performance in low-data regimes (5-10 queries). An in-person user study (N=12) demonstrates that participants can effectively teach their personal contextual preferences with our method, enabling adaptable and personalized reward learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。