arXiv:2507.06224cs.ROcs.AI2025-07ICCV被引 10

无需动作标签视频,让机器人学会处理变形、遮挡等复杂操作。

EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow

  • 基于机器人本体运动学,直接从无标签视频中学习操作动作。
  • 在遮挡、柔性物体和非位移任务上分别提升62%、45%、80%性能。
  • 仅需机器人模型文件即可部署,适合真实场景的通用操作学习。

当前语言引导的机器人操作大多依赖低层级动作标注数据集进行模仿学习。尽管基于物体中心的光流预测方法缓解了这一问题,但仍局限于刚性物体、明显位移且遮挡少的场景。本文提出本体中心光流(EC-Flow)框架,通过预测与本体相关的运动流,直接从无动作标签视频中学习操作。关键洞察在于引入机器人自身的运动学特性,显著提升对多样化操作任务的泛化能力,包括柔性物体处理、遮挡场景及非位移任务。为进一步关联语言指令与物体交互,我们设计目标对齐模块,联合优化运动一致性与目标图像预测。将EC-Flow转化为可执行机器人动作,仅需标准机器人URDF文件定义关节约束,便于实际应用。我们在模拟(Meta-World)和真实世界任务中验证该方法,在遮挡物体处理、柔性物体操作和非位移任务上分别比现有最优物体中心光流方法提升62%、45%、80%。更多信息见项目网站:https://ec-flow1.github.io。

原文摘要 · Abstract (English)

Current language-guided robotic manipulation systems often require low-level action-labeled datasets for imitation learning. While object-centric flow prediction methods mitigate this issue, they remain limited to scenarios involving rigid objects with clear displacement and minimal occlusion. In this work, we present Embodiment-Centric Flow (EC-Flow), a framework that directly learns manipulation from action-unlabeled videos by predicting embodiment-centric flow. Our key insight is that incorporating the embodiment's inherent kinematics significantly enhances generalization to versatile manipulation scenarios, including deformable object handling, occlusions, and non-object-displacement tasks. To connect the EC-Flow with language instructions and object interactions, we further introduce a goal-alignment module by jointly optimizing movement consistency and goal-image prediction. Moreover, translating EC-Flow to executable robot actions only requires a standard robot URDF (Unified Robot Description Format) file to specify kinematic constraints across joints, which makes it easy to use in practice. We validate EC-Flow on both simulation (Meta-World) and real-world tasks, demonstrating its state-of-the-art performance in occluded object handling (62% improvement), deformable object manipulation (45% improvement), and non-object-displacement tasks (80% improvement) than prior state-of-the-art object-centric flow methods. For more information, see our project website at https://ec-flow1.github.io .

机器人操作无监督学习视觉理解动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。