arXiv:2505.16517cs.ROcs.CV2025-05AAAI被引 25

用强化学习让大模型自主学会物理操作,无需人工标注。

ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models

  • 用可验证奖励替代人工标注,实现无监督强化学习。
  • 设计两种规则奖励,提升对操作区域和动作路径的感知能力。
  • 适合研究机器人推理与通用操作的学者快速上手。

大型视觉语言模型(LVLMs)通过视觉感知场景、语言理解指令,推动了机器人操作的发展。然而,现有方法严重依赖昂贵的人工标注数据集,导致泛化能力差,在分布外(OOD)场景中表现不佳,限制了实际应用。为此,我们提出ManipLVM-R1,一种新型强化学习框架,采用可验证奖励(RLVR)替代传统监督学习。该方法直接优化任务对齐结果,提升泛化能力和物理推理,同时摆脱对昂贵标注的依赖。具体而言,我们设计了两种基于规则的奖励函数:作用力感知奖励,用于增强对交互区域的定位;轨迹匹配奖励,确保动作路径的物理合理性。这些奖励提供即时反馈并施加空间逻辑约束,促使模型超越浅层模式匹配,学习更深层次、系统性的物理交互推理。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have recently advanced robotic manipulation by leveraging vision for scene perception and language for instruction following. However, existing methods rely heavily on costly human-annotated training datasets, which limits their generalization and causes them to struggle in out-of-domain (OOD) scenarios, reducing real-world adaptability. To address these challenges, we propose ManipLVM-R1, a novel reinforcement learning framework that replaces traditional supervision with Reinforcement Learning using Verifiable Rewards (RLVR). By directly optimizing for task-aligned outcomes, our method enhances generalization and physical reasoning while removing the dependence on costly annotations. Specifically, we design two rule-based reward functions targeting key robotic manipulation subtasks: an Affordance Perception Reward to enhance localization of interaction regions, and a Trajectory Match Reward to ensure the physical plausibility of action paths. These rewards provide immediate feedback and impose spatial-logical constraints, encouraging the model to go beyond shallow pattern matching and instead learn deeper, more systematic reasoning about physical interactions.

机器人操作强化学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。