arXiv:2502.03356cs.RO2025-02ICRA被引 5

用生成模型反推多人交互中的最优策略,能抗噪声且实时更新。

Inverse Mixed Strategy Games with Generative Trajectory Models

  • 用条件变分自编码器建模混合策略,捕捉多模态行为
  • 在有噪声和目标不确定情况下仍能逼近纳什最优动作
  • 适合需实时适应的机器人协同场景

博弈论模型是建模多智能体交互的有效工具,尤其适用于机器人与人类协作。然而,应用这些模型需要从观测行为中推断其参数,即逆博弈问题,这极具挑战性。现有方法常难以处理行为不确定性与测量噪声,且依赖离线与在线数据。为此,我们提出一种将生成轨迹模型融入可微分混合策略博弈框架的逆博弈方法。通过条件变分自编码器(CVAE)表示混合策略,该方法能在高维、多模态行为分布下,从噪声测量中推理出策略,并实时适应新观测。我们在模拟导航基准上进行了广泛评估,观测由未知博弈模型生成。尽管存在模型失配,我们的方法仍能推导出与真实模型及理想逆博弈基线相当的纳什最优动作,即使在目标不确定和测量噪声下亦然。

原文摘要 · Abstract (English)

Game-theoretic models are effective tools for modeling multi-agent interactions, especially when robots need to coordinate with humans. However, applying these models requires inferring their specifications from observed behaviors -- a challenging task known as the inverse game problem. Existing inverse game approaches often struggle to account for behavioral uncertainty and measurement noise, and leverage both offline and online data. To address these limitations, we propose an inverse game method that integrates a generative trajectory model into a differentiable mixed-strategy game framework. By representing the mixed strategy with a conditional variational autoencoder (CVAE), our method can infer high-dimensional, multi-modal behavior distributions from noisy measurements while adapting in real-time to new observations. We extensively evaluate our method in a simulated navigation benchmark, where the observations are generated by an unknown game model. Despite the model mismatch, our method can infer Nash-optimal actions comparable to those of the ground-truth model and the oracle inverse game baseline, even in the presence of uncertain agent objectives and noisy measurements.

逆博弈生成模型多智能体策略推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。