arXiv:2603.24047cs.RO2026-03

让机器人根据偏好实时调整行为,兼顾速度与节能

PCHC: Enabling Preference Conditioned Humanoid Control via Multi-Objective Reinforcement Learning

  • 用偏好向量控制单一策略,实现多目标动态平衡
  • 在仿真和真实机器人上验证,可实时切换行为优先级
  • 无需训练多个策略,一套模型覆盖多种行为模式

类人机器人常需在速度与能耗等多重目标间权衡。现有强化学习方法虽能掌握复杂技能如防跌倒和感知行走,但受限于固定权重策略,仅能生成单一次优策略,难以提供多样解。本文提出基于多目标强化学习的偏好条件类人控制框架(PCHC),不依赖训练多个策略来逼近帕累托前沿,而是通过单一偏好条件策略实现丰富行为多样性。为此,我们设计基于贝塔分布的对齐机制,利用偏好向量调控混合专家(MoE)模块。在两个代表性类人任务上验证该方法。大量仿真与真实实验表明,该框架可依据输入偏好条件实时调整目标优先级,实现灵活适应。

原文摘要 · Abstract (English)

Humanoid robots often need to balance competing objectives, such as maximizing speed while minimizing energy consumption. While current reinforcement learning (RL) methods can master complex skills like fall recovery and perceptive locomotion, they are constrained by fixed weighting strategies that produce a single suboptimal policy, rather than providing a diverse set of solutions for sophisticated multi-objective control. In this paper, we propose a novel framework leveraging Multi-Objective Reinforcement Learning (MORL) to achieve Preference-Conditioned Humanoid Control (PCHC). Unlike conventional methods that require training a series of policies to approximate the Pareto front, our framework enables a single, preference-conditioned policy to exhibit a wide spectrum of diverse behaviors. To effectively integrate these requirements, we introduce a Beta distribution-based alignment mechanism based on preference vectors modulating a Mixture-of-Experts (MoE) module. We validated our approach on two representative humanoid tasks. Extensive simulations and real-world experiments demonstrate that the proposed framework allows the robot to adaptively shift its objective priorities in real-time based on the input preference condition.

类人控制多目标强化学习偏好条件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。