arXiv:2506.07919cs.LGcs.AI2025-06被引 2

通过渐进线性RNN揭示序列模型中非线性的真正作用

Uncovering the Computational Roles of Nonlinearity in Sequence Modeling Using Almost-Linear RNNs

  • 用几乎线性的RNN逐步减弱非线性,分解网络动态机制
  • 稀疏非线性可提升可解释性、降低计算开销并促进共享表征
  • 在数据少或需离散切换的任务中,稀疏非线性模型表现优于全非线性模型

序列建模任务如自然语言处理、时间序列预测和控制,需要学习复杂的输入输出映射。理论上,非线性递归是实现序列到序列函数通用近似的必要条件,但线性递归模型常表现出意外的有效性。这引出核心问题:非线性何时真正必需?本文提出一种系统框架,剖析递归网络中非线性的功能角色,识别其计算必要性及所支持的机制。方法基于几乎线性递归神经网络(AL-RNNs),允许递归非线性渐进衰减,并将网络动态分解为可分析的线性区间,使计算机制显式化。我们在多种合成与真实世界任务上验证该框架,涵盖经典序列建模基准、神经科学刺激选择任务及多任务套件。结果表明,AL-RNN的分段线性结构能识别出门控、规则集成和依赖记忆的瞬态等计算原语,这些操作在主要线性骨干中自发涌现。稀疏非线性有助于减少并定位非线性计算,提升可解释性,在多任务设置中促进共享表示,同时降低计算成本。此外,稀疏非线性作为有效归纳偏置,在低数据场景或需在不同线性区段间离散切换的任务中,常达到甚至超越全非线性架构的表现。研究为识别非线性功能必要性提供了原则性方法,指导兼具性能、效率与可解释性的递归架构设计。

原文摘要 · Abstract (English)

Sequence modeling tasks across domains such as natural language processing, time series forecasting, and control require learning complex input-output mappings. Nonlinear recurrence is theoretically required for universal approximation of sequence-to-sequence functions, yet linear recurrent models often prove surprisingly effective. This raises the question of when nonlinearity is truly required. We present a framework to systematically dissect the functional role of nonlinearity in recurrent networks, identifying when it is computationally necessary and what mechanisms it enables. We address this using Almost Linear Recurrent Neural Networks (AL-RNNs), which allow recurrence nonlinearity to be gradually attenuated and decompose network dynamics into analyzable linear regimes, making computational mechanisms explicit. We illustrate the framework across diverse synthetic and real-world tasks, including classic sequence modeling benchmarks, a neuroscientific stimulus-selection task, and a multi-task suite. We demonstrate how the AL-RNN's piecewise linear structure enables identification of computational primitives such as gating, rule-based integration, and memory-dependent transients, revealing that these operations emerge within predominantly linear backbones. Across tasks, sparse nonlinearity improves interpretability by reducing and localizing nonlinear computations, promotes shared representations in multi-task settings, and reduces computational cost. Moreover, sparse nonlinearity acts as a useful inductive bias: in low-data regimes or when tasks require discrete switching between linear regimes, sparsely nonlinear models often match or exceed fully nonlinear architectures. Our findings provide a principled approach for identifying where nonlinearity is functionally necessary, guiding the design of recurrent architectures that balance performance, efficiency, and interpretability.

序列建模非线性分析可解释性RNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。