arXiv:2511.05129cs.RO2025-11

分阶段控制抓取与操作,提升机器人对不同物体的泛化能力

Decomposed Object Manipulation via Dual-Actor Policy

  • 采用双角色策略,分别处理接近和操作阶段
  • 在真实场景中平均性能优于当前最优方法10.4%
  • 适用于需要长期多步骤操作的复杂任务

物体操作任务通常可分为接近阶段和操作阶段。以往方法常忽略这一特性,使用单一策略直接学习整个过程。为此,本文提出一种新型双角色策略DAP,显式区分两个阶段,并利用异构视觉先验增强各阶段表现。具体而言,引入基于功能属性的执行者定位操作部件,优化接近过程;再设计基于运动流的执行者捕捉组件运动,促进操作阶段。最后,通过决策模块判断当前阶段并选择对应执行者。此外,现有数据集对象少且缺乏视觉先验支持训练,因此我们构建了包含两种视觉先验的仿真数据集——双先验物体操作数据集,涵盖七项任务,包括两项具有挑战性的长周期、多阶段任务。在该数据集、RoboTwin基准和真实场景中的实验表明,本方法平均性能分别优于当前最优方法5.55%、14.7%和10.4%。

原文摘要 · Abstract (English)

Object manipulation, which focuses on learning to perform tasks on similar parts across different types of objects, can be divided into an approaching stage and a manipulation stage. However, previous works often ignore this characteristic of the task and rely on a single policy to directly learn the whole process of object manipulation. To address this problem, we propose a novel Dual-Actor Policy, termed DAP, which explicitly considers different stages and leverages heterogeneous visual priors to enhance each stage. Specifically, we introduce an affordance-based actor to locate the functional part in the manipulation task, thereby improving the approaching process. Following this, we propose a motion flow-based actor to capture the movement of the component, facilitating the manipulation process. Finally, we introduce a decision maker to determine the current stage of DAP and select the corresponding actor. Moreover, existing object manipulation datasets contain few objects and lack the visual priors needed to support training. To address this, we construct a simulated dataset, the Dual-Prior Object Manipulation Dataset, which combines the two visual priors and includes seven tasks, including two challenging long-term, multi-stage tasks. Experimental results on our dataset, the RoboTwin benchmark and real-world scenarios illustrate that our method consistently outperforms the SOTA method by 5.55%, 14.7% and 10.4% on average respectively.

机器人操作分阶段控制双角色策略仿真数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。