决策能力阈值决定自对弈强化学习是否崩溃。
A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning

- 当所有可触发决策被移除时,智能体迅速陷入固定失败状态。
- 仅保留一个可触发决策点即可避免系统崩溃。
- 适合研究自对弈稳定性与鲁棒性的人参考。
我们发现,决策容量存在一个阈值,决定了自对弈强化学习在规则不对称扰动下是否会崩溃。在扑克变体、矩阵博弈、掷骰子游戏及多种学习算法中,一旦消除所有正可达的可触发决策,智能体将快速收敛至确定性剥削吸引子,即接近最大损失的固定点。即使保留单一正可达的可触发决策点,也能防止此崩溃。冻结基线与固定对手对照实验确认,该机制源于约束下的协同适应,而非扰动本身。该现象与时间无关,动作恢复后完全可逆,并在函数逼近下加剧。结果表明,在测试领域中,零可达加权可触发动作容量处存在明确阈值,严重程度随可达加权容量连续变化。
原文摘要 · Abstract (English)
We show that a threshold in decision capacity determines whether self-play reinforcement learning agents collapse under asymmetric rule perturbations. Across poker variants, matrix games, a dice game, and multiple learning algorithms, eliminating all positive-reach contingent decisions causes rapid convergence to a deterministic exploitation attractor, a fixed point at near-maximal loss. Preserving even a single positive-reach contingent decision point prevents this collapse. A frozen baseline and fixed-opponent control confirm that the mechanism is co-adaptation under constraint, not the perturbation itself. The phenomenon is timing-invariant, fully reversible upon action restoration, and intensifies under function approximation. These results establish a sharp threshold at zero reach-weighted contingent action capacity, with severity scaling continuously via reach-weighted capacity in the tested domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。