arXiv:2606.06194cs.ROcs.CV2026-06被引 4

从第一视角视频中学习主动感知,提升机器人预训练效果

ActiveMimic: Egocentric Video Pretraining with Active Perception

论文配图:ActiveMimic: Egocentric Video Pretraining with Active Perception
图 1 · 摘自论文原文
  • 通过同步恢复相机与手腕轨迹,将镜头运动建模为视角动作
  • 在多种任务中表现超越人类视频预训练基线,媲美机器人数据模型
  • 证明主动感知能力源于人类视频预训练,适合机器人迁移学习

第一人称人类视频为机器人预训练提供了可扩展的数据替代方案,但现有模型在该数据上预训练的效果始终不及机器人数据。我们归因于缺失信号——第一人称视频中的主动感知行为:人类在操作时持续调整视角,导致相机运动,而传统流程将其视为噪声。为此,我们提出ActiveMimic框架,从单个穿戴式RGB摄像头中恢复同步的相机与腕部轨迹,将相机运动建模为视角动作,并在真实世界的第一人称人类视频上联合学习主动感知与操作能力,再迁移到目标机器人。实验证明,在多种具有不同主动感知需求的任务中,ActiveMimic始终优于基于人类视频的基线,且达到机器人数据预训练模型的先进水平。进一步分析表明,主动感知能力源自人类视频预训练,而非机器人微调,证实主动感知是释放第一人称人类视频用于机器人预训练的关键。

原文摘要 · Abstract (English)

Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such video consistently underperform those pretrained on robot data. We attribute this gap to a missing signal, the active perception behavior in egocentric videos, where humans continuously reposition their viewpoint during manipulation, inducing camera motion that standard pipelines treat as noise. To address this, we present ActiveMimic, a pretraining framework that recovers synchronized camera and wrist trajectories from a single body-worn RGB camera, models camera motion as a viewpoint action, and jointly learns active perception and manipulation from in-the-wild egocentric human video before adapting to a target robot. Empirically, real-world experiments across tasks with diverse active perception demands show that ActiveMimic consistently surpasses baselines pretrained on human video and matches state-of-the-art models pretrained on robot data. Further analysis provides evidence that active perception capability originates from egocentric human video pretraining rather than robot-specific fine-tuning, confirming active perception as the key to unlocking egocentric human video for robot pretraining.

第一视角主动感知视频预训练机器人迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。