用流体力学方法高效计算智能体在足球中的接球概率梯度。
Agent Utilities over Generalized Voronoi Regions and their Gradients

- 引入成本诱导的韦洛诺伊区域,将状态空间与划分空间分离。
- 通过积分接球概率密度定义智能体效用,提升计算效率约一个数量级。
- 适用于多智能体博弈场景,尤其适合足球策略优化研究者。
本文推广了韦洛诺伊区域的概念,提出成本诱导的韦洛诺伊(CIV)区域,其中智能体状态空间与划分空间可不同。例如,状态包含位置和速度,而划分仅基于位置。智能体效用定义为在对应CIV区域内对某种效用密度的积分,如足球中接球的概率密度。该效用即为整体接球概率,其梯度可用于优化策略。本文利用流体力学中的雷诺输运定理推导梯度,相比基线有限差分法,在保持相近精度的同时,计算时间减少约一个数量级。
原文摘要 · Abstract (English)
In this paper, we generalize the concept of Voronoi regions, define agent utility as the integral of a utility density over the corresponding Voronoi region, derive gradients of the utility, and illustrate the approach in a two-team example from soccer. The generalization of Voronoi regions is in the form of so-called Cost-Induced Voronoi (CIV) regions, where the agent state space may differ from the space being partitioned. One example of such regions is when the cost is given by the optimal solution of an LQR control problem. Then the agent states include position as well as velocity, while the partitioned space only includes positions. The agent utility is defined by integrating some utility density over the CIV region of the agent. This utility density might be the probability density of some beneficial event, such as receiving a pass in soccer. The utility is then the overall probability of receiving a pass and the gradient represents a way to improve that probability. We show how this utility gradient can be computed using the Reynolds Transport Theorem from fluid mechanics, and that this approach achieves similar accuracy while reducing computation time by about an order of magnitude compared to a baseline finite-difference approximation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。