arXiv:2409.10358eess.AScs.SD2024-09被引 3

对比多种低延迟语音增强技术,用真实数据验证效果。

Ultra-Low Latency Speech Enhancement - A Comprehensive Study

  • 统一训练数据与评估标准,公平比较不同低延迟方法。
  • 实测显示自适应时域滤波器和未来帧预测有效提升性能。
  • 适合听力辅助设备研发者参考,尤其关注实时性场景。

语音增强模型需满足助听设备中低于5毫秒的极低延迟要求。尽管已有多种低延迟技术提出,但使用深度神经网络在受控环境下进行公平比较仍属空白。以往研究在任务、训练数据、脚本和评估设置上存在差异,导致难以进行公正比较。此外,所有方法均在小规模模拟数据集上测试,难以真实反映实际应用表现,可能影响科学结论的可靠性。为此,本文在大规模数据上采用一致训练策略,基于真实世界数据使用更相关指标,全面评估多种低延迟技术的有效性。具体包括非对称窗、可学习窗、自适应时域滤波器组以及未来帧预测技术。同时考察模型增大是否可弥补窗口缩小的影响,并首次探究Mamba架构在低延迟环境中的适用性。

原文摘要 · Abstract (English)

Speech enhancement models should meet very low latency requirements typically smaller than 5 ms for hearing assistive devices. While various low-latency techniques have been proposed, comparing these methods in a controlled setup using DNNs remains blank. Previous papers have variations in task, training data, scripts, and evaluation settings, which make fair comparison impossible. Moreover, all methods are tested on small, simulated datasets, making it difficult to fairly assess their performance in real-world conditions, which could impact the reliability of scientific findings. To address these issues, we comprehensively investigate various low-latency techniques using consistent training on large-scale data and evaluate with more relevant metrics on real-world data. Specifically, we explore the effectiveness of asymmetric windows, learnable windows, adaptive time domain filterbanks, and the future-frame prediction technique. Additionally, we examine whether increasing the model size can compensate for the reduced window size, as well as the novel Mamba architecture in low-latency environments.

语音增强低延迟助听设备Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。