提出软最大反馈系统的精确稳定性阈值,突破传统保守估计。
Sharp Spectral Thresholds for Logit Fixed Points
- 发现无维度的精确稳定条件,取代旧有过度正则化假设。
- 证明在β‖ΠWΠ‖_T→T < 2时系统全局唯一稳定。
- 适用于强化学习、博弈论等需可预测结果的场景。
软最大反馈系统是熵正则化强化学习、对数博弈动态、群体选择和平均场变分更新的核心数学结构。其核心稳定性问题在于:何时自增强的软最大系统会产生唯一且全局可预测的结果?经典理论给出保守答案:将软最大视为单位尺度响应,仅在强随机化条件下保证稳定。我们证明经典方法遗漏了整个稳定区域,未识别出质变真正发生的位置。对于有限维仿射对数系统,其精确的无维度欧氏阈值为 β‖ΠWΠ‖_T→T < 2,而非此前仅在软最大系统高度过正则化时有效的条件。本定理填补了此前缺失的前分岔区域,将仿射软最大反馈系统的稳定性保证扩展至奖励敏感但仍全局可预测的系统,扩大了已认证的稳定边界,并识别出模型真正经历相变的位置。
原文摘要 · Abstract (English)
Softmax feedback systems are a common mathematical core of entropy-regularized reinforcement learning, logit game dynamics, population choice, and mean-field variational updates. Their central stability question is simple: when does a self-reinforcing softmax system produce a unique and globally predictable outcome? Classical theory gives a conservative answer. By treating softmax as a unit-scale response, it certifies stability only in a strongly randomized regime. We prove that the classical approach misses an entire stable regime and does not identify the point at which the qualitative change truly occurs. For finite-dimensional affine logit systems, the sharp dimension-free Euclidean threshold is $$β\|ΠWΠ\|_{\mathcal T\to\mathcal T}<2,$$ rather than the previously used condition, which certifies stability only while the softmax system remains safely over-regularized. Our theorem fills the previously missing pre-bifurcation regime, extending stability guarantees for affine softmax feedback systems to reward-responsive yet globally predictable systems. It enlarges the certified stability boundary for these systems and identifies where the model genuinely undergoes a phase transition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。