arXiv:2606.08934cs.LGstat.AP2026-06

通过反向一致性理论,让循环网络状态更稳定可解释。

Backward Coherence and Hidden-State Stability in Recurrent Neural Networks: A Quasi-Reverse-Martingale Theory

  • 用反向投影重建隐藏状态,构建准逆鞅理论框架
  • 状态稳定性提升43%-58%,收敛速度提前28%-44%
  • 适合需要稳定状态表示的医疗、金融等时序预测场景

循环神经网络维持隐藏状态 $h_t$,但其概率意义常不明确。本文通过研究 extit{反向一致性}:即 $h_t$ 能否通过学习的反向投影 $g_ϕ$ 从 $h_{t+1}$ 重构,揭示隐藏状态的稳定性。在收缩与可求和反向漂移条件下,隐藏状态序列形成准逆鞅,从而获得几乎必然收敛、混合条件下的收敛速率、可解释的极限表示、有限路径停止时间以及时间一致置信序列的理论框架。模拟验证理论:反向一致性正则化使经验准鞅总量 $\ar Q$ 降低43%–58%,稳定性提前28%–44%达成,跟踪误差恢复符合几何界。额外测试确认回声状态遗忘率受 $ρ$ 控制,增量和管 $R_t$ 达到100%同步覆盖(虽偏保守);实践中缺陷尾部代理 $\ar Q_t$ 更有效。反向一致性损失等价于高斯反向模型中的Kullback-Leibler散度,连接变分推断。扩展涵盖 $ϕ$-混合输入、突变点追踪与有限样本浓度。三个真实数据研究进一步验证:在PhysioNet 2012 ICU数据上,逆鞅RNN(RMRNN)匹配传统RNN的死亡预测AUC,且稳定表示提前13小时达成;在FRED-MD上,概念漂移下月度预测误差降低约四倍;在UCI人类活动识别上,过渡后跟踪误差更低并呈几何衰减。保证成立基于所述假设,不宣称普适性。

原文摘要 · Abstract (English)

Recurrent neural networks maintain a hidden state $h_t$, but its probabilistic meaning is often unclear. We study hidden-state stability through \emph{backward coherence}: the extent to which $h_t$ can be reconstructed from $h_{t+1}$ by a learned backward projector $g_ϕ$. Under contraction and summable backward drift, the hidden-state sequence forms a quasi-reverse-martingale. This yields almost-sure convergence, rates under mixing, an interpretable limiting representation, finite pathwise stopping times, and a theoretical framework for time-uniform confidence sequences. Simulations support the theory. Backward-coherence regularisation reduces the empirical quasi-martingale total $\hat Q$ by $43$--$58%$, reaches stability $28$--$44%$ earlier than an unregularised RNN, and gives tracking-error recovery consistent with geometric bounds. Additional tests confirm echo-state forgetting rates bounded by $ρ$ and verify the increment-sum tube $R_t$ with $100%$ simultaneous coverage, although $R_t$ is conservative; in practice, the defect-tail proxy $\hat Q_t$ is the more useful monitor. The backward-coherence loss is also equivalent to minimising a Kullback--Leibler divergence in a Gaussian backward model, linking the method to variational inference. Extensions cover $ϕ$-mixing inputs, change-point tracking, and finite-sample concentration. Three real-data studies further validate the approach. On PhysioNet 2012 ICU data, the Reverse Martingale RNN (RMRNN) matches RNN mortality-prediction AUC while reaching stable representations 13 hours earlier. On FRED-MD, it reduces one-month-ahead forecast error by about fourfold under concept drift. On UCI Human Activity Recognition, it maintains lower post-transition tracking error with geometric decay. The guarantees apply under the stated assumptions; universality is not claimed.

循环网络稳定性理论分析时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。