提出新型延迟记忆单元,提升生物信号分析的时序建模能力
Parallel Delayed Memory Units for Enhanced Temporal Modeling in Biomedical and Bioacoustic Signal Analysis
- 用并行延迟线与勒让德记忆单元压缩时序信息,实现高效状态更新
- 在低信息场景下保持早期特征,显著增强长期记忆能力
- 模块化设计适配多种模型,适合音频与生物信号实时分析
深度学习架构,特别是循环神经网络(RNN),已被广泛应用于音频、生物声学和生物医学信号分析,尤其在数据稀缺环境下表现突出。尽管门控RNN仍具有效性,但在某些情况下参数量过大且训练效率较低;而线性RNN则难以捕捉生物信号中的复杂时序特征。为此,我们提出并行延迟记忆单元(PDMU),一种用于短期时序信用分配的延迟门控状态空间模块,通过门控延迟线机制增强短期时序状态交互与内存效率。不同于以往将时序动态嵌入延迟线结构的延迟记忆单元(DMU),PDMU进一步利用勒让德记忆单元(LMU)将时序信息压缩为向量表示,形成一种因果注意力机制,使模型能动态调整对历史状态的依赖,提升实时学习性能。在低信息场景中,门控机制行为类似跳跃连接,可绕过状态衰减,保留早期表示,从而促进长期记忆。PDMU具备模块化特性,支持并行训练与串行推理,可轻松集成至现有线性RNN框架。此外,我们还引入了双向、高效及脉冲变体,分别在性能或能效上带来额外增益。在多个音频与生物医学基准测试中,实验结果表明PDMU显著提升了记忆容量与整体模型性能。
原文摘要 · Abstract (English)
Advanced deep learning architectures, particularly recurrent neural networks (RNNs), have been widely applied in audio, bioacoustic, and biomedical signal analysis, especially in data-scarce environments. While gated RNNs remain effective, they can be relatively over-parameterised and less training-efficient in some regimes, while linear RNNs tend to fall short in capturing the complexity inherent in bio-signals. To address these challenges, we propose the Parallel Delayed Memory Unit (PDMU), a {delay-gated state-space module for short-term temporal credit assignment} targeting audio and bioacoustic signals, which enhances short-term temporal state interactions and memory efficiency via a gated delay-line mechanism. Unlike previous Delayed Memory Units (DMU) that embed temporal dynamics into the delay-line architecture, the PDMU further compresses temporal information into vector representations using Legendre Memory Units (LMU). This design serves as a form of causal attention, allowing the model to dynamically adjust its reliance on past states and improve real-time learning performance. Notably, in low-information scenarios, the gating mechanism behaves similarly to skip connections by bypassing state decay and preserving early representations, thereby facilitating long-term memory retention. The PDMU is modular, supporting parallel training and sequential inference, and can be easily integrated into existing linear RNN frameworks. Furthermore, we introduce bidirectional, efficient, and spiking variants of the architecture, each offering additional gains in performance or energy efficiency. Experimental results on diverse audio and biomedical benchmarks demonstrate that the PDMU significantly enhances both memory capacity and overall model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。