用分段优势函数提升强化学习中的偏好学习效果
PAWS: Preference Learning with Advantage-Weighted Segments

- 基于分段优势函数直接更新策略,避免步骤级信号误差
- 在模拟机器人任务中性能优于现有偏好强化学习方法
- 适合需要精准时序信用分配的复杂控制场景
基于偏好的强化学习(PbRL)通过人类对轨迹的比较来学习策略,无需显式奖励设计或专家示范。现有方法通常在轨迹或分段层面训练效用函数,但在策略优化时依赖每步的效用估计,导致训练与推理不一致,引发分布偏移,严重影响时序信用分配并限制策略学习。本文分析该问题,提出PAWS——一种基于分段的偏好学习方法,直接使用分段级优势函数进行策略更新。通过使效用训练与策略优化对齐,PAWS保持了轨迹级偏好信息,避免不可靠的每步学习信号。在模拟机器人抓取和行走任务上的实验表明,PAWS始终优于现有PbRL方法,凸显了分布一致性偏好学习的重要性。
原文摘要 · Abstract (English)
Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations. Existing methods typically train utility functions on trajectory or segment-level preferences while relying on per-step utility estimates during policy optimization. This training and inference mismatch induces a distribution shift that severely degrades temporal credit assignment and limits policy learning. We analyze this issue and propose PAWS, a segment-based preference learning method that performs policy updates directly using segment-level advantage functions. By aligning utility training with policy optimization, PAWS preserves trajectory-level preference information and avoids unreliable per-step learning signals. Experiments on simulated robotic manipulation and locomotion tasks demonstrate that PAWS consistently outperforms existing PbRL approaches, highlighting the importance of distribution-consistent preference learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。