arXiv:2606.32027cs.ROcs.AI2026-06被引 1

让机器人通过自然语言偏好学习多维度操作策略,提升长程任务表现。

Freeform Preference Learning for Robotic Manipulation

论文配图:Freeform Preference Learning for Robotic Manipulation
图 1 · 摘自论文原文
  • 用户用自然语言定义偏好轴(如速度、安全),提供多维对比判断。
  • 在6个任务中比稀疏奖励和二元偏好方法提升38个百分点成功率。
  • 支持测试时灵活调整行为方向,无需重新训练,具备行为组合能力。

奖励设计仍是自主机器人策略优化的核心瓶颈,尤其在长程操作任务中,稀疏的成功标签信号不足,而二元偏好又将多种质量概念压缩为单一模糊信号。我们提出自由形式偏好学习(Freeform Preference Learning, FPL),一种从自由语言偏好中学习机器人策略的方法。与要求标注者整体比较两条轨迹不同,FPL允许标注者定义自然语言偏好轴(如速度、安全性、放置质量、谨慎程度),并在每个轴上进行成对偏好判断。这些标注用于训练一个语言条件化的奖励模型,将轨迹与偏好标签映射为特定轴的奖励值。该模型用于训练一个奖励条件策略,在多个人类指定维度上进行优化。在四个真实世界和两个模拟的长程操作任务中,FPL相比稀疏奖励和二元偏好方法提升38个百分点。除性能提升外,FPL无需显式子任务分割即可生成密集进展信号,表现出数据中不存在的行为组合性,并支持测试时无需重训练即可引导策略向不同行为方向演化。

原文摘要 · Abstract (English)

Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of quality into one ambiguous signal. We introduce Freeform Preference Learning (FPL), a method for learning robot policies from freeform human preferences. Rather than asking annotators which of two trajectories is better overall, FPL lets them define natural-language preference axes, such as speed, safety, quality of placement, or carefulness, and provide pairwise preferences along each axis. These annotations are used to learn a language-conditioned reward model that maps a trajectory and preference label to an axis-specific reward. We use this model to train a reward-conditioned policy that optimizes across the multiple human-specified dimensions. Across four real-world and two simulated long-horizon manipulation tasks, FPL improves over sparse-reward and binary-preference methods by 38 percentage points. Beyond improved performance, FPL learns dense progress signals without explicit subtask segmentation, shows compositionality of behavior not present in the data, and allows users to steer the policy towards different behaviors at test time without retraining. Blog post with videos available at https://freeform-pl.github.io/fpl.website/

机器人操控偏好学习多目标优化自然语言指令

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。