arXiv:2608.30326cs.SDcs.AI2026-08中稿 · Interspeech 2026

用并行结构提升语音增强效率,降低误识率。

Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends

论文配图:Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends
图 1 · 摘自论文原文
  • 采用并行时频混合模块,消除递归依赖,提升处理速度
  • 在DNS Challenge和CHiME-4上实现更低的词错误率
  • 无需调参即可抑制语音识别敏感伪影,适合工业部署

语音增强常作为鲁棒语音识别(ASR)的前端,但传统的时序与跨频带模块引入序列依赖,降低并行效率。本文提出一种基于并行时频混合器(PTBM)块的并行分频段增强前端,消除块内递归展开。PTBM在统一并行架构中集成同频时序混合与帧级跨频注意力,实现时间与频率维度的高效上下文建模。系统保留掩码加残差重建接口,并引入可学习的观测添加(LOA),在无需开发集调优的情况下抑制对ASR敏感的伪影。在冻结Whisper后端的DNS Challenge与CHiME-4数据集上,该前端持续降低词错误率,且前端网络仅需0.96 M参数和0.58 GMAC/s。

原文摘要 · Abstract (English)

Speech enhancement is often used as a front-end for robust ASR, yet recurrent temporal and cross-band modules introduce sequential dependencies that reduce parallel efficiency. In this paper, we present a sequence-parallel band-split enhancement front-end built on a Parallel Time-Band Mixer (PTBM) block that eliminates within-block recurrent unrolling. PTBM integrates intra-band temporal mixing and per-frame cross-band attention within a unified parallel architecture, enabling efficient contextual modeling across both time and frequency dimensions. The system retains the mask-plus-residual reconstruction interface and introduces learned Observation-Adding (LOA) to suppress ASR-sensitive artifacts without development-set tuning. Experiments on DNS Challenge and CHiME-4 with frozen Whisper back-ends show that the proposed front-end consistently reduces word error rate relative to recurrent band-split baselines while requiring only 0.96 M parameters and 0.58 GMAC/s for the front-end network.

语音增强并行计算ASR前端模型效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。