受物理能量转换启发,提出新型循环神经网络,长序列建模更高效
Physics-inspired Energy Transition Neural Network for Sequence Learning
- 借鉴物理能量转换模型,设计新型循环结构
- 在多个序列任务上超越Transformer,且计算复杂度更低
- 适合追求低开销长序列建模的场景
近年来,Transformer因其卓越性能成为比传统循环神经网络(RNN)更稳健、可扩展的序列建模方案。然而,Transformer在捕捉长期依赖方面的能力主要源于其全面的成对建模过程,而非内在的序列语义归纳偏置。本文重新探索纯RNN的能力,评估其长期学习机制。受物理能量转换模型启发,我们提出一种名为「物理启发的能量转换神经网络」(PETNN)的有效循环结构。实验表明,PETNN的记忆机制能有效存储长期依赖信息,在多个序列任务上优于基于Transformer的方法。此外,由于其循环特性,PETNN具有显著更低的计算复杂度。本研究提出了一种最优的基础循环架构,揭示了在当前由Transformer主导的领域中发展高效循环神经网络的潜力。
原文摘要 · Abstract (English)
Recently, the superior performance of Transformers has made them a more robust and scalable solution for sequence modeling than traditional recurrent neural networks (RNNs). However, the effectiveness of Transformer in capturing long-term dependencies is primarily attributed to their comprehensive pair-modeling process rather than inherent inductive biases toward sequence semantics. In this study, we explore the capabilities of pure RNNs and reassess their long-term learning mechanisms. Inspired by the physics energy transition models that track energy changes over time, we propose a effective recurrent structure called the``Physics-inspired Energy Transition Neural Network" (PETNN). We demonstrate that PETNN's memory mechanism effectively stores information over long-term dependencies. Experimental results indicate that PETNN outperforms transformer-based methods across various sequence tasks. Furthermore, owing to its recurrent nature, PETNN exhibits significantly lower complexity. Our study presents an optimal foundational recurrent architecture and highlights the potential for developing effective recurrent neural networks in fields currently dominated by Transformer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。