解决远场语音增强中自声回响问题,实现低延迟清晰语音输出。
Don't Listen to Me: A Lightweight, Low-Latency Model for Own-Voice Cancellation in Far-Field Speech Enhancement
- 用2毫秒低延迟模型,基于短语音样本实现目标人声消除。
- 新方法在信噪比和主观评分上优于现有模型,计算量更低。
- 适合智能音箱等需要实时语音反馈的远场设备使用。
我们提出自声消除(OVC):在嘈杂多说话人混合语音中移除目标说话人(注册用户)的声音,同时保留其他语音内容。这与目标说话人提取互为补充,旨在解决远场设备将增强音频回传给用户时因延迟过高导致的自声失真问题——往返延迟常超过人耳感知阈值。我们采用仅需2毫秒算法延迟的时间域模型,仅依赖一段短注册语音进行条件建模,并对比了基于TD-SpeakerBeam的方法与更轻量的Mamba-MinGRU掩码器(由Mamba块与MinGRU时间混合构成)。通过将原有基于ConvTasNet的辅助网络替换为线性RNN编码器,不仅提升了信号失真比(SDR)与预测的MOS分,还显著降低了计算开销。结果表明,OVC是远场降噪中一项实用且低延迟的优化目标。
原文摘要 · Abstract (English)
We introduce own-voice cancellation (OVC): removing a target (enrolled) speaker from a noisy multi-speaker mixture while preserving any remaining speech. Framed as the complement of target speaker extraction, OVC addresses latency-induced own-voice artifacts that arise when a far-field device streams enhanced audio back to the user, as the round-trip time easily exceeds the perceptual threshold for own-voice distortion. We condition a time-domain model with only 2 ms algorithmic latency on a short enrollment utterance and benchmark TD-SpeakerBeam alongside a lighter Mamba-MinGRU masker built from Mamba blocks with MinGRU temporal mixing. Replacing the ConvTasNet-based auxiliary network with a linear RNN encoder improves both signal-to-distortion ratio and predicted MOS while reducing compute. Results establish OVC as a practical, low-latency enhancement objective for far-field denoising.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。