arXiv:2505.13094cs.SDcs.AI2025-05被引 2

用注意力缓存记忆提升实时语音分离的长期信息捕捉能力

Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech Separation

  • 引入时频注意力缓存机制,通过记忆模块存储历史信息
  • 在相同性能下参数量更少,计算复杂度显著降低
  • 适合部署于资源受限的实时语音处理场景

现有因果语音分离模型因难以保留历史信息而性能不及非因果模型。为此,本文提出时频注意力缓存记忆(TFACM)模型,通过注意力机制与缓存记忆(CM)有效捕捉时空关系。模型中,LSTM层建模频率相对位置,时间维度采用局部与全局表示实现因果建模。CM模块存储过往信息,因果注意力精炼(CAR)模块进一步优化时间特征表示,提升粒度。实验表明,TFACM性能接近当前最优的TF-GridNet-Causal模型,但参数量和计算复杂度大幅降低。

原文摘要 · Abstract (English)

Existing causal speech separation models often underperform compared to non-causal models due to difficulties in retaining historical information. To address this, we propose the Time-Frequency Attention Cache Memory (TFACM) model, which effectively captures spatio-temporal relationships through an attention mechanism and cache memory (CM) for historical information storage. In TFACM, an LSTM layer captures frequency-relative positions, while causal modeling is applied to the time dimension using local and global representations. The CM module stores past information, and the causal attention refinement (CAR) module further enhances time-based feature representations for finer granularity. Experimental results showed that TFACM achieveed comparable performance to the SOTA TF-GridNet-Causal model, with significantly lower complexity and fewer trainable parameters. For more details, visit the project page: https://cslikai.cn/TFACM/.

语音分离注意力机制缓存记忆实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。