研究多智能体环境中智能体持续学习时的性能退化问题,提出可预测失败的关键结构
Reinforcing the World's Edge: A Continual Learning Problem in the Multi-Agent-World Boundary
- 定义'不变核心'捕捉成功轨迹中的高覆盖率抽象模式
- 证明轨迹分布漂移最多使成功率下降ε/p₀,且系数不可改进
- 实验验证核心退化能提前预警失败,适用于持续控制与迁移学习场景
在静态去中心化马尔可夫博弈中,学习型同伴为焦点智能体生成随回合变化的诱导MDP序列。尽管联合博弈保持静态,焦点智能体的奖励与动态仍发生漂移,形成以智能体为中心的持续强化学习问题。在回合内固定同伴策略可保留焦点轨迹的分布与期望回报。然而,成功条件下的可复用结构可能因同伴策略更新而退化。本文提出‘不变核心’概念,通过在多数成功轨迹中出现的高频率抽象模式表征该结构。主要结果为最坏情况紧致的条件定理:轨迹分布漂移ε最多使候选项的成功条件覆盖率下降至ε/p₀,其中p₀为其参考成功质量,该系数是紧的。同伴策略移动提供ε;正覆盖率余量由此可得Ω(1/η)的认证生存周期,并在显式有效冲突条件下(如解析类中精确策略梯度实现),达到匹配的Θ(1/η)首次退出律。经校准的成功质量与可执行性下,同一证书亦可导出策略价值、库选择与迁移后悔保证。一个可精确求解的走廊环境验证了结构预测,包括反比例寿命关系(R²>0.9999)。两个注册的64流研究(持续控制与cue-MNIST)表明,核心退化可预测即将失败并实现近似最优干预;对八个已学伙伴的基于层级觅食发展配对的探索性再分析也支持核心退化与失败之间的关联。
原文摘要 · Abstract (English)
In a stationary decentralized Markov game, learning peers generate an episode-indexed sequence of induced MDPs for any focal agent. The joint game remains stationary while the focal agent's rewards and dynamics drift, forming an agent-centric continual reinforcement-learning problem. Marginalizing peers whose policies are fixed within an episode preserves every focal trajectory law and expected return. Success-conditioned reusable structure may therefore degrade under peer updates. An \emph{invariant core} represents such structure through maximal abstract patterns appearing in a high fraction of successful focal trajectories. The main result is a worst-case-tight conditioning theorem: trajectory-law drift $\varepsilon$ can reduce a candidate's success-conditioned coverage by at most $\frac{\varepsilon}{p_0}$, where $p_0$ is its reference success mass, and the coefficient is sharp. Peer-policy movement supplies $\varepsilon$; positive coverage margin then yields a certified $Ω(\frac{1}η)$ survival horizon and, under an explicit effective-conflict condition realized by exact policy gradient in an analytic class, a matching $Θ(\frac{1}η)$ first-exit law. With calibrated success mass and executability, the same certificate yields policy-value, library-selection, and transfer-regret guarantees. An exactly solvable corridor confirms the structural predictions, including the inverse-rate lifetime ($R^2>0.9999$). Two registered 64-stream studies in continual control and cue-MNIST show that core erosion predicts impending failure and enables near-oracle intervention; an exploratory reanalysis of eight learned-partner Level-Based Foraging development pairings suggests the same erosion--failure link under peer learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。