arXiv:2606.07012cs.RO2026-06被引 1

通过拆解任务组件实现多样化演示生成,提升机器人3D操作的泛化能力。

Task Editing for Generalizable 3D Visuomotor Policy Learning

论文配图:Task Editing for Generalizable 3D Visuomotor Policy Learning
图 1 · 摘自论文原文
  • 将任务分解为场景、技能、物体三部分,灵活重组生成新轨迹。
  • 在真实机器人上验证,显著提升长周期操作任务的性能与泛化性。
  • 适用于难采集场景,如障碍物避让和复杂杂乱环境。

3D视觉运动策略为复杂机器人操作提供了有前景的方向,因其深度图和点云可提供丰富的几何信息用于空间推理。然而,其成功通常依赖大规模真实世界示范,而收集这些数据成本高且耗时。现有方法常通过对象中心变换(如改变物体姿态或尺度)生成示范以提高数据效率,但此类变换大多保持原始场景结构与技能序列,难以合成多样化的场景-技能-物体组合,限制了对复杂任务的泛化能力。本文提出任务编辑框架 Task-Edit,从任务中心视角生成多样化轨迹。核心思想是将任务分解为场景、技能和物体三个组件,并灵活重组。该方法实现了可扩展的示范生成,显著提升了长周期操作任务的泛化能力。我们在大量真实世界实验中评估了 Task-Edit,验证了三大优势:(1) 有效性:在多种真实任务和机器人形态下显著提升3D视觉运动策略性能;(2) 泛化性:模型在不同场景设置下表现出更强适应能力;(3) 可适用性:使模型能处理现实中难以采集的场景,包括抗干扰、避障及未见过的杂乱场景。

原文摘要 · Abstract (English)

3D visuomotor policies offer a promising direction for complex robotic manipulation, as depth maps and point clouds provide rich geometric information for spatial reasoning. However, their success often depends on large-scale real-world demonstrations, which are costly and time-consuming to collect. To this end, existing methods commonly use demonstration generation strategies to improve data efficiency by applying object-centric transformations to human-collected demonstrations, such as varying object poses or scales. While effective for local variation, these transformations largely preserve the original scene structure and skill sequence, limiting their ability to synthesize diverse scene-skill-object combinations for complex tasks. In this paper, we propose Task-Edit, a novel demonstration generation framework that generates diverse trajectories from a task-centric editing perspective. The key insight of Task-Edit is to decompose a task into scene, skill and object components, and flexibly recombine them. In this way, Task-Edit enables scalable demonstration generation and significantly improves generalization for long-horizon manipulation tasks. We evaluate Task-Edit through extensive real-world experiments and demonstrate three advantages: (1) Effectiveness: Task-Edit significantly improves 3D visuomotor policies across various real-world tasks and robot embodiments. (2) Generalizability: Task-Edit improves model generalization across different scenario setups. (3) Applicability: Task-Edit enables models to handle scenarios that are difficult to collect in the real world, including disturbance resistance, obstacle avoidance and unseen cluttered scenes.

机器人操作3D视觉泛化能力演示生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。