arXiv:2509.19626cs.ROcs.CV2025-09NeurIPS被引 30

让机器人从第一视角人类数据中通用模仿,跨域对齐提升成功率44%。

EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data

  • 用最优传输对齐人机策略隐空间,保留动作相关特征。
  • 在3个真实任务上比基线提升44%成功率,且泛化到新物体新场景。
  • 适合做机器人模仿学习、跨设备迁移的开发者参考。

第一人称人类经验数据为端到端机器人操作模仿学习提供了巨大资源。然而,人类与机器人在视觉外观、传感器模态和运动学上的显著域差距阻碍了知识迁移。本文提出EgoBridge,一种统一的协同训练框架,通过域适应显式对齐人类与机器人数据的策略隐空间。基于最优传输(OT)计算联合策略隐特征与动作之间的差异度量,学习出不仅在人类与机器人域间对齐,且保留对策略学习至关重要的动作相关信息的观测表示。EgoBridge在三个真实世界的单臂和双臂操作任务中,相比人类增强的跨体基线,政策成功率达绝对提升44%。该方法还能泛化至仅在人类数据中出现的新物体、新场景和新任务,而基线完全失效。视频与更多详情见https://ego-bridge.github.io

原文摘要 · Abstract (English)

Egocentric human experience data presents a vast resource for scaling up end-to-end imitation learning for robotic manipulation. However, significant domain gaps in visual appearance, sensor modalities, and kinematics between human and robot impede knowledge transfer. This paper presents EgoBridge, a unified co-training framework that explicitly aligns the policy latent spaces between human and robot data using domain adaptation. Through a measure of discrepancy on the joint policy latent features and actions based on Optimal Transport (OT), we learn observation representations that not only align between the human and robot domain but also preserve the action-relevant information critical for policy learning. EgoBridge achieves a significant absolute policy success rate improvement by 44% over human-augmented cross-embodiment baselines in three real-world single-arm and bimanual manipulation tasks. EgoBridge also generalizes to new objects, scenes, and tasks seen only in human data, where baselines fail entirely. Videos and additional information can be found at https://ego-bridge.github.io

模仿学习域适应机器人第一人称数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。