提出嵌套训练框架NestRL,让AI更好适应会变的人类队友。
NestRL: A Nested Training Regime for Mutual Adaptation in Human-AI Teaming
- 分层训练:每层AI对抗下一层的自适应代理,模拟人类动态调整
- 实测在Overcooked中比基线提升任务完成率,对陌生伙伴和真人表现更优
- 理论与实验均证明能避免死板协作策略,适合人机协同场景
人机协同中的相互适应是核心挑战,因人类会根据AI行为调整策略。现有方法通过多样化训练伙伴近似人类行为,但这些伙伴通常静态,无法捕捉人类的动态性。标准多智能体联合训练常导致代理收敛到仅与特定伙伴有效的隐晦协作策略,泛化能力差。为此,我们把人机协同建模为交互式部分可观马尔可夫决策过程(I-POMDP),提出嵌套训练框架NestRL:在每一层训练代理时,对抗来自下一层的自适应代理,使其暴露于动态行为同时避免生成不透明协作策略。理论分析表明,NestRL代理不会收敛到依赖特定伙伴的策略;在Overcooked环境中实证验证,其在面对未见过的自适应代理及真实人类队友时,任务性能更高,且交互过程中展现出显著更强的适应能力。
原文摘要 · Abstract (English)
Mutual adaptation is a central challenge in human-AI teaming, as humans naturally adjust their strategies in response to an AI agent's behavior. Existing approaches attempt to approximate human behavior by diversifying training partners; however, these partners are typically static and fail to capture the adaptive nature of human teammates. When agents are trained jointly in standard multi-agent settings, they often converge to opaque coordination strategies that work only with their co-trained partners, leading to poor generalization. To model adaptive human behavior, we formulate human-AI teaming as an Interactive Partially Observable Markov Decision Process (I-POMDP). We propose NestRL, a nested training regime that learns the solution to a finite-level I-POMDP by training agents at each level against adaptive agents from the level below. This exposes agents to adaptive behavior while preventing emergence of opaque coordination strategies. We provide theoretical analysis showing that NestRL agents avoid convergence to partner-specific strategies, and validate this empirically in the Overcooked domain against state-of-the-art baselines. NestRL achieves higher task performance with both unseen adaptive agents and real human teammates, while exhibiting significantly greater adaptability over the course of interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。