arXiv:2502.02377cs.AI2025-02中稿 · AAMAS 2025被引 4

针对临时协作中的伙伴不确定性,提出对抗性优化策略提升鲁棒性。

A Minimax Approach to Ad Hoc Teamwork

  • 用对抗先验替代固定分布,显式建模伙伴不确定性
  • 在Melting Pot烹饪任务中优于自对弈、虚构博弈等方法
  • 适合需要应对未知合作者的现实协作场景

我们提出一种最小最大-贝叶斯方法来解决临时协作(AHT)问题,该方法在部署时针对伙伴的对抗性先验优化策略,显式处理对伙伴的不确定性。与现有假设特定伙伴分布的方法不同,本方法提升了最坏情况下的性能保障。大量实验,包括在Melting Pot套件中的协同烹饪任务评估,表明该方法在鲁棒性上优于自对弈、虚构博弈和最优响应学习。研究强调了选择合适的队友训练分布对实现AHT鲁棒性的关键作用。

原文摘要 · Abstract (English)

We propose a minimax-Bayes approach to Ad Hoc Teamwork (AHT) that optimizes policies against an adversarial prior over partners, explicitly accounting for uncertainty about partners at time of deployment. Unlike existing methods that assume a specific distribution over partners, our approach improves worst-case performance guarantees. Extensive experiments, including evaluations on coordinated cooking tasks from the Melting Pot suite, show our method's superior robustness compared to self-play, fictitious play, and best response learning. Our work highlights the importance of selecting an appropriate training distribution over teammates to achieve robustness in AHT.

临时协作鲁棒性对抗优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。