arXiv:2509.24784cs.LGcs.AI2025-09NeurIPS被引 1

新基准Labyrinth可精准测试模仿学习的泛化能力。

Quantifying Generalisation in Imitation Learning

  • 构建可控环境,分离训练、评估与测试场景
  • 支持可重复实验,验证不同泛化因素的影响
  • 适合研究泛化性与鲁棒性的人工智能学者

模仿学习基准常因训练与评估设置差异不足,难以有效评估泛化性能。我们提出Labyrinth,一个可精确控制结构、起始与目标位置及任务复杂度的基准环境,实现训练、评估与测试设置的明确区分。该环境具备离散且完全可观测的状态空间,以及已知最优动作,支持可解释性与细粒度评估。其灵活架构可针对性测试泛化因素,包含部分可观测性、钥匙-门任务与冰面障碍等变体。通过支持受控、可复现的实验,Labyrinth推动了模仿学习泛化评估的发展,为开发更鲁棒智能体提供重要工具。

原文摘要 · Abstract (English)

Imitation learning benchmarks often lack sufficient variation between training and evaluation, limiting meaningful generalisation assessment. We introduce Labyrinth, a benchmarking environment designed to test generalisation with precise control over structure, start and goal positions, and task complexity. It enables verifiably distinct training, evaluation, and test settings. Labyrinth provides a discrete, fully observable state space and known optimal actions, supporting interpretability and fine-grained evaluation. Its flexible setup allows targeted testing of generalisation factors and includes variants like partial observability, key-and-door tasks, and ice-floor hazards. By enabling controlled, reproducible experiments, Labyrinth advances the evaluation of generalisation in imitation learning and provides a valuable tool for developing more robust agents.

模仿学习泛化评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。