arXiv:2607.28737cs.LGcs.CV2026-07

用第三人称观察学习模仿,让模型像镜子一样理解他人动作。

Mirror Learning

论文配图:Mirror Learning
图 1 · 摘自论文原文
  • 通过视频扩散模型转换视角,生成第一人称模拟数据
  • 仅用镜像数据即可训练出有效策略,比传统方法更优
  • 适合缺乏真人示范数据的机器人学习场景

我们从第三人称观察的角度研究模仿学习,提出镜像学习框架:从被动观察中获取可执行策略。虽然行为克隆(BC)在密集且对齐的第一人称数据下表现优异,但无法利用人类和动物常用来学习的丰富第三人称演示信号。我们提出一种方法,包含(i)使用微调的视频扩散模型实现视角转换,使学习者‘换位思考’;(ii)逆动力学模型推断学习者控制空间中的动作轨迹。这使得能合成镜像数据——由第三人称观察生成的伪第一人称专家数据。实验证明,仅使用镜像数据即可训练出有效策略,且将镜像数据用于第一人称BC训练能进一步提升下游策略性能。结果表明,现代生成世界模型隐式包含足够结构,可提供一种可扩展且安全的替代方案,减少对遥操作数据采集的依赖。

原文摘要 · Abstract (English)

We investigate imitation learning through the lens of third-person observation and propose a framework for mirror learning: acquiring actionable policies from passive observation. While behavior cloning (BC) excels under dense, well-aligned first-person data, it fundamentally fails to leverage the rich observational signals arising from third-person demonstrations that humans and animals routinely exploit. We introduce a method that composes (i) a learned perspective transformation that places learners in demonstrators' shoes using a fine-tuned video diffusion model and (ii) an inverse dynamics model that infers action trajectories in the learners' control space. This enables the synthesis of mirror data, pseudo first-person expert data generated from third-person observations of demonstrator behavior. Empirically, we show that mirror data alone can train effective policies, and that augmenting first-person BC training with mirror data further improves downstream policy performance. Our results suggest that modern generative world models implicitly encode sufficient structure to enable a scalable and safe alternative to teleoperation-heavy data collection.

模仿学习视频生成镜像学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。