arXiv:2603.02426cs.LG2026-03

多智能体协作学习共享结构,提升个性化策略收敛速度。

Personalized Multi-Agent Average Reward TD-Learning via Joint Linear Approximation

  • 通过联合线性近似分解共享子空间与局部权重
  • 实现线性加速,有效缓解信号冲突问题
  • 适合需要协同学习的异构多智能体场景

我们研究个性化多智能体平均奖励TD学习,其中多个智能体在不同环境中交互并共同学习各自的值函数。重点考虑存在共享线性表示且最优权重整体位于未知线性子空间的情形。受个性化联邦学习成功启发,我们分析了合作式单时标TD学习的收敛性,即智能体迭代估计公共子空间与局部头。研究表明,该分解可过滤冲突信号,有效缓解‘错位’信号的负面影响,并实现线性加速。主要技术挑战源于异质性、马尔可夫采样及其复杂交互对误差演化的影响:多个变量的误差动态紧密耦合,且最优子空间与估计子空间间的主角距离无直接收缩性。我们希望分析方法能为深入挖掘共性结构提供启示。实验验证了通过共享结构学习对更一般控制问题的益处。

原文摘要 · Abstract (English)

We study personalized multi-agent average reward TD learning, in which a collection of agents interacts with different environments and jointly learns their respective value functions. We focus on the setting where there exists a shared linear representation, and the agents' optimal weights collectively lie in an unknown linear subspace. Inspired by the recent success of personalized federated learning (PFL), we study the convergence of cooperative single-timescale TD learning in which agents iteratively estimate the common subspace and local heads. We showed that this decomposition can filter out conflicting signals, effectively mitigating the negative impacts of ``misaligned'' signals, and achieving linear speedup. The main technical challenges lie in the heterogeneity, the Markovian sampling, and their intricate interplay in shaping error evolutions. Specifically, not only are the error dynamics of multiple variables closely interconnected, but there is also no direct contraction for the principal angle distance between the optimal subspace and the estimated subspace. We hope our analytical techniques can be useful to inspire research on deeper exploration into leveraging common structures. Experiments are provided to show the benefits of learning via a shared structure to the more general control problem.

多智能体TD学习共享结构线性加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。