让机器人通过双向互动学习空间感知与动作生成
X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction

- 构建双向循环框架,动作与姿态相互修正
- 24个仿真和3个真实任务中表现优于现有方法
- 适合需要精细操作的机器人系统研发者
有效处理空间感知与动作生成之间的交互仍是机器人操作中的关键瓶颈。现有方法通常将空间感知与动作执行视为解耦或单向过程,从根本上限制了机器人掌握复杂操作任务的能力。为此,我们提出 X-Imitator,一个通用的双路径框架,将空间感知与动作执行建模为紧密耦合的双向循环。通过相互条件化当前姿态预测与过去动作,该框架实现了空间推理与动作生成的持续互馈优化,精准模拟人类内部前向模型。系统采用模块化设计,可无缝集成至多种视觉运动策略中。在24个仿真任务和3个真实世界任务上的大量实验表明,本框架显著优于基线策略及使用显式姿态引导的先前方法。代码将开源。
原文摘要 · Abstract (English)
Effectively handling the interplay between spatial perception and action generation remains a critical bottleneck in robotic manipulation. Existing methods typically treat spatial perception and action execution as decoupled or strictly unidirectional processes, fundamentally restricting a robot's ability to master complex manipulation tasks. To address this, we propose X-Imitator, a versatile dual-path framework that models spatial perception and action execution as a tightly coupled bidirectional loop. By reciprocally conditioning current pose predictions on past actions and vice versa, this framework enables continuous mutual refinement between spatial reasoning and action generation. This joint modeling exactly mimics human internal forward models. Designed as a modular architecture, the system can be seamlessly integrated into various visuomotor policies. Extensive experiments across 24 simulated and 3 real-world tasks demonstrate that our framework significantly outperforms both vanilla policies and prior methods utilizing explicit pose guidance. The code will be open sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。