arXiv:2602.09043eess.AScs.LG2026-02中稿 · ICASSP 2026, Barce…

用局部+全局摘要提升低资源语音识别效率,节省40%显存。

Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition

  • 窗口化摘要混合:融合局部邻域与全局摘要,保持线性复杂度。
  • 仅微调替换块,使显存降低40%,推理延迟更少。
  • 适合资源受限场景,如边缘设备语音识别任务。

自监督学习(SSL)虽推动了语音处理发展,但因自注意力机制存在二次时间复杂度。为解决此问题,已有研究提出线性时间的摘要混合(SM)方法,通过均值池化总结整个语音片段,但缺乏足够的局部上下文。本文提出窗口化摘要混合(WSM),在保留线性复杂度的同时,引入局部邻域摘要以增强时间依赖性。同时,提出选择性微调策略:将SSL模型中的自注意力层替换为WSM块,并仅微调这些块。该方法在低资源语音识别中显著提升性能,同时使峰值显存使用量减少40%。WSM块具备线性时间复杂度和更强的上下文感知能力。选择性替换部分注意力层可降低计算、内存与延迟,适用于低资源语音识别场景。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has advanced speech processing but suffers from quadratic complexity due to self-attention. To address this, SummaryMixing (SM) has been proposed as a linear-time alternative that summarizes entire utterances using mean pooling but lacks sufficient local context. In this work, we introduce Windowed SummaryMixing (WSM), which enhances SM by integrating local neighborhood summaries alongside the global summary, maintaining efficiency while improving temporal dependencies. Additionally, we introduce a selective fine-tuning approach, replacing self-attention layers in SSL models with WSM blocks and fine-tuning only these blocks in low-resource settings. Our approach improves ASR performance while reducing peak VRAM usage by 40\% in the SSL models. WSM blocks have linear-time complexity with enhanced context awareness. Selectively replacing some attention layers reduces compute, memory, and latency, making it ideal for low-resource speech recognition.

语音识别自监督学习高效模型低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。