arXiv:2605.24343cs.AI2026-05

让AI学会根据搭档特点灵活调整动作,提升人机协作效率。

Adaptive Human-AI Coordination via Hierarchical Action Disentanglement

论文配图:Adaptive Human-AI Coordination via Hierarchical Action Disentanglement
图 1 · 摘自论文原文
  • 分层强化学习框架,用高层技能控制底层动作序列
  • 在多种布局和未见过的伙伴下,表现优于现有方法
  • 适合需要动态适应的实时人机协作场景

人机协作需要能适应不同合作方行为和能力水平的智能体,同时对未知合作者保持鲁棒性。现有方法常退化为单一主导行为或学习到不匹配的技能,限制了有效协作。我们提出内在动作解耦(IAD),一种深度分层强化学习(DHRL)框架,学习与高层潜在技能相关的、面向合作方的低层动作序列。IAD引入内在奖励,显式鼓励低层策略在不同技能间产生解耦的动作分布,实现高层决策与特定合作方行为响应之间的可解释映射。通过捕捉时间上延展的交互模式,IAD在分布外变化下仍能灵活适应异构合作动态。我们在Overcooked-AI环境中评估IAD,覆盖多个布局及多样化合作设置,包括未见过的模拟合作者、基于人类-人类游戏数据训练的人类代理模型以及真实人类合作者。结果表明,IAD在所有设置中均持续优于强基线,展现出更可靠、自适应的协作能力。

原文摘要 · Abstract (English)

Human-AI collaboration requires agents that can adapt to diverse partner behaviors and skill levels while remaining robust to unseen partners. Existing methods often collapse to a single dominant behavior or learn poorly aligned skills, limiting effective coordination. We propose Intrinsic Action Disentanglement (IAD), a deep hierarchical reinforcement learning (DHRL) framework that learns distinct, partner-aware low-level action sequences conditioned on high-level latent skills. IAD introduces an intrinsic reward that explicitly encourages disentangled action distributions of the agent's low-level policy across skills, yielding an interpretable mapping between high-level decisions and partner-specific behavioral responses. By capturing temporally extended interaction patterns, IAD enables flexible adaptation to heterogeneous partner dynamics under distributional shift. We evaluate IAD in the Overcooked-AI domain across multiple layouts and diverse partner settings, including unseen simulated partners, a human-proxy model trained on human-human gameplay, and real human partners. Results show that IAD consistently outperforms strong baselines and achieves more reliable, adaptive coordination across all settings.

人机协作分层强化学习动作解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。