arXiv:2411.19309cs.ROcs.CV2024-11被引 111

让机器人模型通过偏好对齐,更好应对未知任务和安全效率需求。

GRAPE: Generalizing Robot Policy via Preference Alignment

论文配图:GRAPE: Generalizing Robot Policy via Preference Alignment
图 1 · 摘自论文原文
  • 用成功与失败轨迹对齐,提升模型泛化能力
  • 在真实和仿真环境中,成功率提升超50%
  • 可灵活适配安全、高效等不同目标

尽管视觉-语言-动作(VLA)模型在多种机器人任务中取得进展,但其仅依赖成功轨迹进行行为克隆,导致在未见任务上泛化能力差。此外,模型通常在不同设置下微调专家示范数据,引入分布偏差,限制对效率、安全、任务完成等多样目标的适应性。为此,我们提出GRAPE:基于偏好对齐的通用机器人策略。GRAPE在轨迹层面对齐VLA,并从成功与失败尝试中隐式建模奖励,以增强对多样化任务的泛化能力。同时,它将复杂操作任务分解为独立阶段,利用大视觉-语言模型提出的关键点,自动施加定制化的时空约束,引导偏好建模。这些约束灵活可调,可对齐安全、效率或任务成功等目标。我们在真实与仿真环境中评估GRAPE,结果表明其显著提升现有最优VLA模型性能,在域内和未见操作任务上的成功率分别提高51.79%和58.20%。此外,通过偏好对齐,碰撞率降低37.44%,轨迹步长减少11.15%。所有代码、模型与数据已公开于https://grape-vla.github.io/

原文摘要 · Abstract (English)

Despite the recent advancements of vision-language-action (VLA) models on a variety of robotics tasks, they suffer from critical issues such as poor generalizability to unseen tasks, due to their reliance on behavior cloning exclusively from successful rollouts. Furthermore, they are typically fine-tuned to replicate demonstrations collected by experts under different settings, thus introducing distribution bias and limiting their adaptability to diverse manipulation objectives, such as efficiency, safety, and task completion. To bridge this gap, we introduce GRAPE: Generalizing Robot Policy via Preference Alignment. Specifically, GRAPE aligns VLAs on a trajectory level and implicitly models reward from both successful and failure trials to boost generalizability to diverse tasks. Moreover, GRAPE breaks down complex manipulation tasks to independent stages and automatically guides preference modeling through customized spatiotemporal constraints with keypoints proposed by a large vision-language model. Notably, these constraints are flexible and can be customized to align the model with varying objectives, such as safety, efficiency, or task success. We evaluate GRAPE across a diverse array of tasks in both real-world and simulated environments. Experimental results demonstrate that GRAPE enhances the performance of state-of-the-art VLA models, increasing success rates on in-domain and unseen manipulation tasks by 51.79% and 58.20%, respectively. Additionally, GRAPE can be aligned with various objectives, such as safety and efficiency, reducing collision rates by 37.44% and rollout step-length by 11.15%, respectively. All code, models, and data are available at https://grape-vla.github.io/

机器人控制偏好对齐泛化能力多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。