arXiv:2409.14542cs.LGcs.MA2024-09

通过鲁棒方法从观测信号反推多智能体感知系统的效用函数。

Distributionally Robust Inverse Reinforcement Learning for Identifying Multi-Agent Coordinated Sensing

  • 构建基于Wasserstein集的最坏情况优化框架,提升效用估计鲁棒性。
  • 证明了鲁棒估计与半无限优化等价,确保解的理论一致性。
  • 适用于认知雷达等存在噪声观测的协同感知系统建模。

我们推导了一种最小-最大分布鲁棒逆强化学习(IRL)算法,用于重建多智能体感知系统的效用函数。具体地,我们构造了在以噪声信号观测为中心的Wasserstein模糊集上最小化最坏情况预测误差的效用估计器。我们证明了这种鲁棒估计与一个半无限优化重构问题等价,并提出了一个一致的算法来计算解。我们在数值实验中展示了该鲁棒IRL方案的有效性,成功从观测跟踪信号重构了认知雷达网络的效用函数。

原文摘要 · Abstract (English)

We derive a minimax distributionally robust inverse reinforcement learning (IRL) algorithm to reconstruct the utility functions of a multi-agent sensing system. Specifically, we construct utility estimators which minimize the worst-case prediction error over a Wasserstein ambiguity set centered at noisy signal observations. We prove the equivalence between this robust estimation and a semi-infinite optimization reformulation, and we propose a consistent algorithm to compute solutions. We illustrate the efficacy of this robust IRL scheme in numerical studies to reconstruct the utility functions of a cognitive radar network from observed tracking signals.

逆强化学习多智能体鲁棒优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。