研究多智能体在线控制中对抗扰动下的鲁棒性,提出基于梯度的局部更新算法。
Online Multi-Agent Control with Adversarial Disturbances
- 采用在线梯度法,每智能体本地更新策略以应对对抗性扰动。
- 证明了个体后悔值为次线性且接近最优,与智能体数量有不同依赖关系。
- 适用于目标一致的场景,可跟踪动态博弈均衡,适合复杂系统实时控制。
在线多智能体控制问题广泛存在于自动驾驶、经济系统和能源管理等领域,其中各智能体追求随时间变化的竞争目标。在这些场景中,对对抗性扰动的鲁棒性至关重要。本文研究受对抗性扰动影响的多智能体线性动力系统中的在线控制问题。不同于以往大多假设无噪声或随机扰动的研究,本文考虑对抗性扰动下的在线设定,每个智能体试图最小化其自身的凸损失序列。在两种反馈模型下,分析基于梯度的控制器及其局部策略更新,证明了关于时间跨度的次线性且近似最优的个体后悔界,并揭示了不同智能体数量下的缩放特性。当智能体目标一致时,多智能体控制问题转化为时变势博弈,进而推导出均衡跟踪保证。结果首次将在线控制与在线学习博弈相衔接,为动态连续状态环境中的个体与集体性能提供了鲁棒保障。
原文摘要 · Abstract (English)
Online multi-agent control problems, where many agents pursue competing and time-varying objectives, are widespread in domains such as autonomous robotics, economics, and energy systems. In these settings, robustness to adversarial disturbances is critical. In this paper, we study online control in multi-agent linear dynamical systems subject to such disturbances. In contrast to most prior work in multi-agent control, which typically assumes noiseless or stochastically perturbed dynamics, we consider an online setting where disturbances can be adversarial, and where each agent seeks to minimize its own sequence of convex losses. Under two feedback models, we analyze online gradient-based controllers with local policy updates. We prove per-agent regret bounds that are sublinear and near-optimal in the time horizon and that highlight different scalings with the number of agents. When agents' objectives are aligned, we further show that the multi-agent control problem induces a time-varying potential game for which we derive equilibrium tracking guarantees. Together, our results take a first step in bridging online control with online learning in games, establishing robust individual and collective performance guarantees in dynamic continuous-state environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。