让AI学会根据搭档行为调整策略,提升人机协作稳定性。
Partner-Aware Hierarchical Skill Discovery for Robust Human-AI Collaboration

- 设计新框架让技能学习依赖搭档行为,避免盲目套用固定模式。
- 在过厨游戏测试中,对不同水平和风格的伙伴表现更优,胜过现有方法。
- 适合需要灵活适应人类或异质智能体的协作场景,如机器人陪练、虚拟助手。
多智能体协作,尤其是人机协同,要求智能体能适应具有多样性和动态行为的新搭档。传统深度分层强化学习(DHRL)方法聚焦于以智能体为中心的奖励,忽视搭档行为,导致捷径学习——即技能利用虚假信息而非真正适应搭档的动态行为,削弱了智能体的适应与协调能力。本文提出伙伴感知型分层技能发现(PASD),一种基于伙伴行为条件的DHRL框架。PASD引入对比性内在奖励,捕捉由伙伴交互产生的模式,在相似伙伴间对齐技能表示,同时保持不同策略间的可区分性。通过基于伙伴交互结构化技能空间,该方法缓解了捷径学习问题,促进行为一致性,实现稳健且可适应的协调。我们在包含多样化伙伴(不同技能水平和玩法风格)的Overcooked-AI基准上进行广泛评估,并使用从人-人游戏轨迹训练的人类代理模型进一步验证。结果表明,PASD始终优于现有群体基和分层基线,展现出跨多种伙伴行为的可迁移技能学习能力。对学习到的技能表示分析显示,PASD能有效适配多样化伙伴行为,凸显其在人机协作中的鲁棒性。
原文摘要 · Abstract (English)
Multi-agent collaboration, especially in human-AI teaming, requires agents that can adapt to novel partners with diverse and dynamic behaviors. Conventional Deep Hierarchical Reinforcement Learning (DHRL) methods focus on agent-centric rewards and overlook partner behavior, leading to shortcut learning, where skills exploit spurious information instead of adapting to partners' dynamic behaviors. This limitation undermines agents' ability to adapt and coordinate effectively with novel partners. We introduce Partner-Aware Skill Discovery (PASD), a DHRL framework that learns skills conditioned on partner behavior. PASD introduces a contrastive intrinsic reward to capture patterns emerging from partner interactions, aligning skill representations across similar partners while maintaining discriminability across diverse strategies. By structuring the skill space based on partner interactions, this approach mitigates shortcut learning and promotes behavioral consistency, enabling robust and adaptive coordination. We extensively evaluate PASD in the Overcooked-AI benchmark with a diverse population of partners characterized by varying skill levels and play styles. We further evaluate the approach with human proxy models trained from human-human gameplay trajectories. PASD consistently outperforms existing population-based and hierarchical baselines, demonstrating transferable skill learning that generalizes across a wide range of partner behaviors. Analysis of learned skill representations shows that PASD adapts effectively to diverse partner behaviors, highlighting its robustness in human-AI collaboration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。