arXiv:2603.21091stat.MLcs.LG2026-03
研究非马尔可夫环境下随机逼近的新框架,解析Transformer与持续学习的理论基础。
Stochastic approximation in non-markovian environments revisited
- 构建非遍历非马尔可夫过程下的随机逼近分析框架
- 揭示注意力机制与持续学习依赖完整历史的内在机理
- 为自回归模型和长时依赖学习提供理论支撑
基于作者近期关于非马尔可夫环境中随机逼近的研究,本文进一步探讨了驱动随机过程既非马尔可夫又非遍历的情况。在此基础上,提出一个分析框架,用于理解基于Transformer的学习机制,特别是注意力机制,以及持续学习。这两类方法在原则上都依赖于完整的过去信息。该框架为分析具有长期依赖性的序列建模提供了新的理论视角。
原文摘要 · Abstract (English)
Based on some recent work of the author on stochastic approximation in non-markovian environments, the situation when the driving random process is non-ergodic in addition to being non-markovian is considered. Using this, we propose an analytic framework for understanding transformer based learning, specifically, the `attention' mechanism, and continual learning, both of which depend on the entire past in principle.
随机逼近Transformer持续学习非马尔可夫
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。