测试ICRL在临时团队协作中的极限,发现现有方法表现不佳。
Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

- 构建大规模基准ICRL4AHT,支持多智能体即时适应评估
- 百万级交互数据表明,现有方法在未知队友下常低于随机基线
- 揭示部分可观测环境下的策略推理挑战,适合多智能体协同研究者
上下文强化学习(ICRL)使基础智能体能快速适应新任务,但在临时团队协作(AHT)——需与未知伙伴协作——场景下的有效性尚未被探索。为此,我们提出大规模基准ICRL4AHT,基于高吞吐的JAX实现的Overcooked-V2。该基准包含多样化的队友集合,涵盖强化学习与启发式策略,支持受控的训练-测试分布偏移,并提供可复现的端到端流程:队友生成、学习历史收集、数据集构建与在线多轮评估。我们在数百万次交互中评估代表性历史条件型ICRL算法,如算法蒸馏(AD)和决策预训练变换器(DPT)。结果揭示显著局限:尽管在单智能体任务中成功,这些基线在多智能体场景中无法实现稳健的测试时适应。具体而言,它们在未见队友和未见布局任务中频繁劣于随机基线,且在长时程下无明显上下文改进。这些发现凸显了在OvercookedV2 AHT协议下部分可观测环境中的策略推理挑战,确立本基准为下一代协调算法的关键测试平台。
原文摘要 · Abstract (English)
In-Context Reinforcement Learning (ICRL) has enabled foundation agents to adapt instantaneously to novel tasks, yet its efficacy in Ad-Hoc Teamwork (AHT)-where coordination with unknown partners is required-remains unexplored. To rigorously evaluate this, we introduce a large-scale benchmark ICRL4AHT, built upon a high-throughput JAX implementation of Overcooked-V2. Our benchmark includes a large, diverse teammate suite spanning both RL and heuristic policies, enabling controlled train-test shifts, and provides a reproducible end-to-end pipeline for teammate generation, learning-history collection, dataset construction, and online multi-episode evaluation. We evaluate representative history-conditioned ICRL algorithms, including Algorithm Distillation (AD) and Decision-Pretrained Transformer (DPT), across millions of transitions. Results reveal notable limitations: contrary to their success in single-agent domains, these baselines fail to exhibit robust test-time adaptation in multi-agent settings. Specifically, these methods frequently underperform random baselines across both unseen teammate and unseen layout tracks, with no clear in-context improvement over long horizons. These findings highlight the challenges of strategic inference under partial observability within the OvercookedV2 AHT protocol, establishing our benchmark as a critical testbed for next-generation coordination algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。