AMaze生成可交互的迷宫基准,提升智能体泛化能力。
AMaze: An intuitive benchmark generator for fast prototyping of generalizable agents
- 通过视觉符号设计复杂迷宫,支持人类参与生成与策略分析。
- 交互式训练使泛化性能提升50%至100%,最优效果达100%增益。
- 适合研究人机协同训练、泛化能力评估的科研人员使用。
传统智能体训练多基于单一、确定性的简单环境,但此类环境缺乏泛化能力。近期基准常采用多环境设置,如随机噪声或环境变换,但大多依赖人工设计或随机生成,成本高且难控。本文提出AMaze,一种新型基准生成器,要求具身智能体通过解读任意复杂度和欺骗性视觉符号来导航迷宫。该生成器支持特征定制迷宫生成,便于人类理解智能体策略。以离散场景为例,对比三种训练方式(一次性、渐进式、交互式),结果表明后两者在泛化能力上显著优于直接训练:根据评估指标、训练方式与算法组合,中位性能提升50%至100%,交互式训练表现最佳,验证了可控人机协同生成器的价值。
原文摘要 · Abstract (English)
Traditional approaches to training agents have generally involved a single, deterministic environment of minimal complexity to solve various tasks such as robot locomotion or computer vision. However, agents trained in static environments lack generalization capabilities, limiting their potential in broader scenarios. Thus, recent benchmarks frequently rely on multiple environments, for instance, by providing stochastic noise, simple permutations, or altogether different settings. In practice, such collections result mainly from costly human-designed processes or the liberal use of random number generators. In this work, we introduce AMaze, a novel benchmark generator in which embodied agents must navigate a maze by interpreting visual signs of arbitrary complexities and deceptiveness. This generator promotes human interaction through the easy generation of feature-specific mazes and an intuitive understanding of the resulting agents' strategies. As a proof-of-concept, we demonstrate the capabilities of the generator in a simple, fully discrete case with limited deceptiveness. Agents were trained under three different regimes (one-shot, scaffolding, interactive), and the results showed that the latter two cases outperform direct training in terms of generalization capabilities. Indeed, depending on the combination of generalization metric, training regime, and algorithm, the median gain ranged from 50% to 100% and maximal performance was achieved through interactive training, thereby demonstrating the benefits of a controllable human-in-the-loop benchmark generator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。