arXiv:2502.09287cs.LG2025-02被引 2

揭示线性循环网络在长序列建模中的性能极限与权衡规律。

An Uncertainty Principle for Linear Recurrent Neural Networks

  • 通过经典信号模型分析线性RNN的滤波能力,构建近似移位K步的滤波器。
  • 证明最优滤波器需在第K步前平均值的范围正比于K/S,实现理论下界。
  • 揭示信息处理中精度与宽度的不确定性关系,适合研究序列建模原理者阅读。

我们研究线性递归神经网络,这类网络因其在长程建模中具备稳定且高效的能力,已成为序列建模的核心组件。本文聚焦于一个基础但关键的复制任务:构建一个阶数为S的线性滤波器,以逼近一个向过去看K步的移位-K滤波器(shift-K filter),其中K > S。基于经典信号模型与二次代价函数,我们完整刻画了该问题,给出了逼近误差的下界,并构造出能达到该下界(常数因子内)的显式滤波器。最优性能揭示了一个不确定性原理:最优滤波器必须对过去第K步附近的值进行平均,其平均范围(宽度)与K/S成正比。

原文摘要 · Abstract (English)

We consider linear recurrent neural networks, which have become a key building block of sequence modeling due to their ability for stable and effective long-range modeling. In this paper, we aim at characterizing this ability on a simple but core copy task, whose goal is to build a linear filter of order $S$ that approximates the filter that looks $K$ time steps in the past (which we refer to as the shift-$K$ filter), where $K$ is larger than $S$. Using classical signal models and quadratic cost, we fully characterize the problem by providing lower bounds of approximation, as well as explicit filters that achieve this lower bound up to constants. The optimal performance highlights an uncertainty principle: the optimal filter has to average values around the $K$-th time step in the past with a range~(width) that is proportional to $K/S$.

序列建模线性RNN不确定性原理滤波器设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。