通过共享结构学习实现无需微调的仿真到现实迁移
BIFROST: Bridging Invariant Feature Representation for Observation-space Sim2Real Transfer

- 基于跨域双模拟目标学习共享状态编码器
- 在视觉与动力学差异下实现零样本迁移
- 适合需要高泛化能力的机器人操控任务
仿真到现实迁移在机器人策略学习中受限于仿真与现实之间的差异。现有方法通常分别处理各差异,但缺乏对任务在仿真与现实间共享结构的利用。本文提出BIFROST,通过配对跨域数据,使用跨域双模拟目标学习共享历史编码器:无论领域差异如何,产生等效长期行为的观测-动作序列均被映射至相近潜在状态。在仿真中训练的策略可直接迁移到现实,实现零样本转移。在仿真到仿真视觉导航、仿真到现实接触丰富操作与视觉伺服任务中,实验表明BIFROST在存在视觉与动力学差异时,优于领域适应与联合训练基线。
原文摘要 · Abstract (English)
Sim2real transfer for robot policy learning suffers due to mismatch between simulation and reality. Existing methods typically address each gap in isolation through separate adaptation modules, which are composed or layered when both gaps coexist. Yet the basis for attempting sim2real in the first place is that there is shared structure between a task in simulation and reality, where equivalent actions from equivalent configurations produce equivalent long term outcomes regardless of domain specific differences in rendering or physics. In this paper, we study whether we can identify and exploit this shared structure from raw observations to train a policy that enables zero shot transfer. We introduce BIFROST, which learns a shared history encoder on paired cross-domain data via cross-domain bisimulation objective: observation-action sequences leading to equivalent long-term behavior are mapped to nearby latent states, regardless of domain. Policies trained on these latent states in simulation transfer zero-shot to reality. We provide empirical evidence on sim2sim visual navigation and sim2real contact rich manipulation task and visual servoing task that BIFROST achieves effective transfer where domain adaptation and co-training baselines fail under both visual and dynamics domain gaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。