arXiv:2409.13683cs.RO2024-09被引 4

用多模态变换器建模人类偏好,提升机器人行为对齐效果

PrefMMT: Modeling Human Preferences in Preference-based Reinforcement Learning with Multimodal Transformers

  • 分离状态与动作模态,用分层变换器捕捉时序与交互关系
  • 在D4RL和Meta-World任务上优于现有最优基线
  • 适合关注人机对齐与复杂行为建模的研究者

基于人类偏好的强化学习(PbRL)在对齐机器人行为与人类偏好方面展现出潜力,但其成效高度依赖于奖励模型对人类偏好的准确建模。现有方法多采用马尔可夫假设进行偏好建模(PM),忽略了机器人轨迹中影响人类评估的时序依赖性。尽管近期工作通过序列建模学习非马尔可夫奖励以缓解此问题,却忽视了机器人轨迹的多模态特性——由状态与动作两种不同模态构成。因此,它们难以捕捉两种模态间复杂的相互作用,而这正是塑造人类偏好的关键。本文提出一种多模态序列建模方法用于偏好建模,通过解耦状态与动作模态,引入名为PrefMMT的多模态变换器网络,层次化利用模态内时序依赖与模态间状态-动作交互,以捕捉复杂的偏好模式。实验表明,PrefMMT在D4RL基准的行走任务与Meta-World基准的操控任务上均持续优于当前最优的偏好建模基线。

原文摘要 · Abstract (English)

Preference-based reinforcement learning (PbRL) shows promise in aligning robot behaviors with human preferences, but its success depends heavily on the accurate modeling of human preferences through reward models. Most methods adopt Markovian assumptions for preference modeling (PM), which overlook the temporal dependencies within robot behavior trajectories that impact human evaluations. While recent works have utilized sequence modeling to mitigate this by learning sequential non-Markovian rewards, they ignore the multimodal nature of robot trajectories, which consist of elements from two distinctive modalities: state and action. As a result, they often struggle to capture the complex interplay between these modalities that significantly shapes human preferences. In this paper, we propose a multimodal sequence modeling approach for PM by disentangling state and action modalities. We introduce a multimodal transformer network, named PrefMMT, which hierarchically leverages intra-modal temporal dependencies and inter-modal state-action interactions to capture complex preference patterns. We demonstrate that PrefMMT consistently outperforms state-of-the-art PM baselines on locomotion tasks from the D4RL benchmark and manipulation tasks from the Meta-World benchmark.

偏好建模多模态强化学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。