通过主动探索提升参数识别,让腿式机器人仿真到现实的迁移更精准。
Sampling-Based System Identification with Active Exploration for Legged Robot Sim2Real Learning
- 基于大规模并行采样估计机器人物理参数,减少仿真与真实轨迹差异。
- 主动探索策略优化输入命令,使采集数据更具信息量,提升参数辨识精度。
- 在多种运动任务中比基线方法性能高42%-63%,适合复杂足式机器人部署。
仿真到现实的差异限制了学习型策略在真实世界中实现高精度任务。尽管领域随机化(DR)常被用于弥合这一差距,但其通常依赖启发式方法,且调参不当会导致策略过于保守、性能下降。系统辨识(Sys-ID)提供了针对性方案,但传统方法依赖可微分动力学和直接力矩测量,这在接触丰富的足式系统中很少成立。为此,我们提出SPI-Active(基于采样的参数识别与主动探索),一个两阶段框架,通过大规模并行采样鲁棒地估计足式机器人的关键物理参数,最小化仿真与真实轨迹之间的状态预测误差。为进一步提高数据的信息量,我们引入主动探索策略,通过优化探索策略的输入指令,最大化所收集真实轨迹的费雪信息量。该定向探索显著提升了参数辨识准确性,并增强了跨多样化任务的泛化能力。实验表明,SPI-Active使学习策略在真实世界中的仿真到现实迁移更加精确,在多种运动任务中性能优于基线42%-63%。
原文摘要 · Abstract (English)
Sim-to-real discrepancies hinder learning-based policies from achieving high-precision tasks in the real world. While Domain Randomization (DR) is commonly used to bridge this gap, it often relies on heuristics and can lead to overly conservative policies with degrading performance when not properly tuned. System Identification (Sys-ID) offers a targeted approach, but standard techniques rely on differentiable dynamics and/or direct torque measurement, assumptions that rarely hold for contact-rich legged systems. To this end, we present SPI-Active (Sampling-based Parameter Identification with Active Exploration), a two-stage framework that estimates physical parameters of legged robots to minimize the sim-to-real gap. SPI-Active robustly identifies key physical parameters through massive parallel sampling, minimizing state prediction errors between simulated and real-world trajectories. To further improve the informativeness of collected data, we introduce an active exploration strategy that maximizes the Fisher Information of the collected real-world trajectories via optimizing the input commands of an exploration policy. This targeted exploration leads to accurate identification and better generalization across diverse tasks. Experiments demonstrate that SPI-Active enables precise sim-to-real transfer of learned policies to the real world, outperforming baselines by 42-63% in various locomotion tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。