arXiv:2502.12330cs.ROcs.LG2025-02被引 10

X-IL框架系统探索模仿学习设计空间,发现更优策略配置。

X-IL: Exploring the Design Space of Imitation Learning Policies

  • 模块化设计支持灵活替换骨干网络与优化方法。
  • 新配置在机器人学习基准上性能显著超越现有方法。
  • 适合想系统优化模仿学习的科研与工程人员。

设计现代模仿学习(IL)策略需权衡特征编码、架构、策略表示等多个决策。随着领域快速演进,可用选项持续增多,导致IL策略的设计空间庞大且尚未充分探索。本文提出X-IL——一个开源可访问的框架,用于系统性探索该设计空间。其模块化结构支持无缝替换策略组件,如骨干网络(如Transformer、Mamba、xLSTM)和优化技术(如Score-matching、Flow-matching)。这种灵活性促进了全面实验,并发现了若干优于现有方法的新策略配置,在近期机器人学习基准上表现突出。实验不仅展示了显著性能提升,还揭示了各类设计选择的优劣。本研究既为实践者提供参考,也为未来模仿学习研究奠定基础。

原文摘要 · Abstract (English)

Designing modern imitation learning (IL) policies requires making numerous decisions, including the selection of feature encoding, architecture, policy representation, and more. As the field rapidly advances, the range of available options continues to grow, creating a vast and largely unexplored design space for IL policies. In this work, we present X-IL, an accessible open-source framework designed to systematically explore this design space. The framework's modular design enables seamless swapping of policy components, such as backbones (e.g., Transformer, Mamba, xLSTM) and policy optimization techniques (e.g., Score-matching, Flow-matching). This flexibility facilitates comprehensive experimentation and has led to the discovery of novel policy configurations that outperform existing methods on recent robot learning benchmarks. Our experiments demonstrate not only significant performance gains but also provide valuable insights into the strengths and weaknesses of various design choices. This study serves as both a practical reference for practitioners and a foundation for guiding future research in imitation learning.

模仿学习策略优化框架设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。