将控制与感知分离训练,用少量真实数据实现高效机器人操作
Best of Sim and Real: Decoupled Visuomotor Manipulation via Learning Control in Simulation and Perception in Real
- 控制在仿真中训练,感知仅在真实环境部署时适配
- 仅需10-20次真实演示即达成优秀性能
- 适合追求低实测数据需求的机器人系统开发者
由于端到端学习中感知与控制耦合,仿真到现实的迁移仍是机器人操作的核心挑战。本文提出一种解耦框架:控制策略在仿真中利用特权状态训练,掌握空间布局与操作动力学;感知则仅在部署时适应真实观测,以对齐冻结的控制策略。核心洞察是控制策略和动作模式具有跨环境通用性,可通过系统性随机化在仿真中学习;而感知具固有域特定性,必须在真实视觉观测下学习。相比需大量真实数据的端到端方法,本方法仅需10-20次真实演示即可实现优异表现,将复杂的仿真到现实问题简化为结构化的感知对齐任务。我们在桌面操作任务上验证该方法,结果显示其具备更强的数据效率与分布外泛化能力,能有效应对超出训练分布的物体位置与尺度,证实解耦感知与控制可显著提升仿真到现实的迁移效果。
原文摘要 · Abstract (English)
Sim-to-real transfer remains a fundamental challenge in robot manipulation due to the entanglement of perception and control in end-to-end learning. We present a decoupled framework that learns each component where it is most reliable: control policies are trained in simulation with privileged state to master spatial layouts and manipulation dynamics, while perception is adapted only at deployment to bridge real observations to the frozen control policy. Our key insight is that control strategies and action patterns are universal across environments and can be learned in simulation through systematic randomization, while perception is inherently domain-specific and must be learned where visual observations are authentic. Unlike existing end-to-end approaches that require extensive real-world data, our method achieves strong performance with only 10-20 real demonstrations by reducing the complex sim-to-real problem to a structured perception alignment task. We validate our approach on tabletop manipulation tasks, demonstrating superior data efficiency and out-of-distribution generalization compared to end-to-end baselines. The learned policies successfully handle object positions and scales beyond the training distribution, confirming that decoupling perception from control fundamentally improves sim-to-real transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。