arXiv:2606.12328eess.AS2026-06中稿 · Interspeech 2026

HALO通过降半帧率减少语音增强模型冗余计算,提升效率。

HALO: Half-Frame-Rate Adaptive Learnable Operator for Lightweight STFT-Based Speech Enhancement

论文配图:HALO: Half-Frame-Rate Adaptive Learnable Operator for Lightweight STFT-Based Speech Enhancement
图 1 · 摘自论文原文
  • 在不改变STFT流程前提下,用自适应卷积降低内部帧率
  • 在DNS3数据集上,轻量模型性能统一提升且无延迟增加
  • 适合资源受限场景下的实时语音增强系统

基于STFT的语音增强通常采用重叠分析帧,虽保障了稳定处理,但相邻帧高度相关,导致轻量模型出现冗余计算。本文提出因果可插拔模块HALO,不改变STFT过程,将内部帧率减半。HALO在骨干网络前进行自适应降帧率,在恢复阶段重建原始帧率谱图,两者均使用轻量动态卷积实现。帧率减半有效降低主干计算开销,且无算法延迟增加,释放资源用于拓宽通道。在DNS3数据集上的实验表明,在相同复杂度下,多种轻量模型均获得一致性能提升,验证了消除重叠冗余的有效性。

原文摘要 · Abstract (English)

STFT-based speech enhancement typically adopts overlapping analysis frames. While overlap is essential for stable STFT processing, it makes adjacent frames highly correlated, causing redundant computation in lightweight models. We propose Half-frame-rate Adaptive Learnable Operator (HALO), a causal plug-in module that halves the internal frame rate without altering the STFT procedure. Broadly applicable to many lightweight models, HALO applies adaptive rate reduction before the backbone and restoration afterward, reconstructing the full-rate spectrum on the original STFT grid. Both reduction and restoration are implemented with lightweight dynamic convolutions. By halving the processed frame rate, HALO reduces backbone compute cost with no added algorithmic latency, freeing budget for channel widening. Experiments on the DNS3 dataset show consistent gains across diverse lightweight models under matched complexity, demonstrating the effectiveness of reducing overlap-induced redundancy.

语音增强轻量模型STFT高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。