arXiv:2412.07544cs.LGcs.RO2024-12ICLR被引 6

用收缩动力系统提升机器人在未知环境中的恢复能力

Contractive Dynamical Imitation Policies for Efficient Out-of-Sample Recovery

  • 设计可收缩的动力系统策略,保证所有轨迹收敛
  • 在模拟机械臂与导航任务中显著提升未知区域表现
  • 适合需要高可靠性的机器人控制场景

模仿学习通过专家行为数据训练策略,但在分布外(OOS)区域常表现不可靠。现有基于稳定动力系统的方法虽能保证收敛,却忽略瞬态行为。本文提出基于收缩动力系统的策略框架,确保任意参数下策略轨迹均收敛,从而实现高效分布外恢复。利用循环平衡网络与耦合层构建策略结构,保证对任意参数选择均满足收缩性,支持无约束优化。同时提供最坏情况与期望损失的理论上限,严格验证部署可靠性。实验表明,在模拟机器人抓取与导航任务中,该方法显著提升分布外性能。

原文摘要 · Abstract (English)

Imitation learning is a data-driven approach to learning policies from expert behavior, but it is prone to unreliable outcomes in out-of-sample (OOS) regions. While previous research relying on stable dynamical systems guarantees convergence to a desired state, it often overlooks transient behavior. We propose a framework for learning policies modeled by contractive dynamical systems, ensuring that all policy rollouts converge regardless of perturbations, and in turn, enable efficient OOS recovery. By leveraging recurrent equilibrium networks and coupling layers, the policy structure guarantees contractivity for any parameter choice, which facilitates unconstrained optimization. We also provide theoretical upper bounds for worst-case and expected loss to rigorously establish the reliability of our method in deployment. Empirically, we demonstrate substantial OOS performance improvements for simulated robotic manipulation and navigation tasks.

模仿学习机器人控制动力系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。