让智能体通过共享任务抽象,实现零样本协作。
Implicitly Aligning Humans and Autonomous Agents through Shared Task Abstractions
- 用分层强化学习模拟人类协作的结构化思维
- 在Overcooked中对新队友和环境变化表现更稳定
- 适合需要快速适应陌生伙伴的协同场景
在协作任务中,自主智能体难以像人类一样快速适应新队友。我们认为,零样本协作能力受限于缺乏共享任务抽象这一人类依赖的隐式对齐机制。为此,我们提出HA²:分层即兴智能体框架,利用分层强化学习模仿人类协作的结构化方式。我们在Overcooked环境中评估了HA²,结果表明其在与未见过的智能体及人类协作时,相比现有基线有统计学显著提升,对环境变化更具鲁棒性,且优于所有前沿方法。
原文摘要 · Abstract (English)
In collaborative tasks, autonomous agents fall short of humans in their capability to quickly adapt to new and unfamiliar teammates. We posit that a limiting factor for zero-shot coordination is the lack of shared task abstractions, a mechanism humans rely on to implicitly align with teammates. To address this gap, we introduce HA$^2$: Hierarchical Ad Hoc Agents, a framework leveraging hierarchical reinforcement learning to mimic the structured approach humans use in collaboration. We evaluate HA$^2$ in the Overcooked environment, demonstrating statistically significant improvement over existing baselines when paired with both unseen agents and humans, providing better resilience to environmental shifts, and outperforming all state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。