arXiv:2601.08122cs.LG2026-01被引 1

提出新方法分析RNN的泛化能力,提升其在分布外数据上的鲁棒性。

Generalization Analysis and Method for Domain Generalization for a Family of Recurrent Neural Networks

  • 用科尔曼算子将RNN状态演化建模为线性系统,实现可解释性。
  • 通过谱分析量化领域偏移对泛化误差的最大影响。
  • 新方法显著降低分布外误差,适合时序数据安全应用。

深度学习在科学与工程领域推动了广泛进展,但其模型常表现出有限的可解释性和泛化能力,尤其在安全关键场景中可能削弱信任。因此,研究者日益关注(i)分析可解释性与泛化能力,以及(ii)开发在训练数据分布之外仍表现稳健的模型(即领域泛化)。然而,深度学习的理论分析仍不完整,例如多数泛化分析假设样本独立,这在具有时间相关性的序列数据中不成立。针对这些局限,本文提出一种分析方法,用于一类循环神经网络(RNNs)的可解释性与域外(OOD)泛化。具体而言,将训练后RNN的状态演化视为未知的离散时间非线性闭环反馈系统,利用科尔曼算子理论将其近似为线性算子,从而实现可解释性。随后采用谱分析量化领域偏移对泛化误差的最坏影响。基于此分析,提出一种领域泛化方法,可降低域外泛化误差并增强对分布偏移的鲁棒性。最后,该分析与泛化方法在实际的时间模式学习任务上得到验证。

原文摘要 · Abstract (English)

Deep learning (DL) has driven broad advances across scientific and engineering domains. Despite its success, DL models often exhibit limited interpretability and generalization, which can undermine trust, especially in safety-critical deployments. As a result, there is growing interest in (i) analyzing interpretability and generalization and (ii) developing models that perform robustly under data distributions different from those seen during training (i.e. domain generalization). However, the theoretical analysis of DL remains incomplete. For example, many generalization analyses assume independent samples, which is violated in sequential data with temporal correlations. Motivated by these limitations, this paper proposes a method to analyze interpretability and out-of-domain (OOD) generalization for a family of recurrent neural networks (RNNs). Specifically, the evolution of a trained RNN's states is modeled as an unknown, discrete-time, nonlinear closed-loop feedback system. Using Koopman operator theory, these nonlinear dynamics are approximated with a linear operator, enabling interpretability. Spectral analysis is then used to quantify the worst-case impact of domain shifts on the generalization error. Building on this analysis, a domain generalization method is proposed that reduces the OOD generalization error and improves the robustness to distribution shifts. Finally, the proposed analysis and domain generalization approach are validated on practical temporal pattern-learning tasks.

RNN泛化分析领域泛化可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。