提出新基准评估语音编解码器在复杂环境下的重建与任务一致性
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
- 构建环境鲁棒语音编解码器评测基准ERSB,系统评估抗干扰能力
- 复杂环境使重建质量与下游任务一致性显著下降,峰值下降超20%
- 揭示现有编解码器在真实场景中的局限性,适合语音处理研究者参考
神经语音编解码器在纯净语音重建上表现优异,但在复杂声学环境及下游信号处理任务中的有效性仍缺乏深入研究。本文提出新型基准环境鲁棒语音编解码器基准(ERSB),系统评估编解码器的环境鲁棒性。具体评估两个关键能力:(1) 鲁棒重建,衡量语音与非语音声学细节的保留程度;(2) 下游任务一致性,确保使用重建语音替代原始语音时,下游信号处理任务偏差最小。全面实验表明,复杂声学环境显著降低信号重建质量与下游任务一致性。该工作揭示了当前语音编解码器的局限性,并指明提升其环境鲁棒性的未来方向。
原文摘要 · Abstract (English)
Neural speech codecs excel in reconstructing clean speech signals; however, their efficacy in complex acoustic environments and downstream signal processing tasks remains underexplored. In this study, we introduce a novel benchmark named Environment-Resilient Speech Codec Benchmark (ERSB) to systematically evaluate whether neural speech codecs are environment-resilient. Specifically, we assess two key capabilities: (1) robust reconstruction, which measures the preservation of both speech and non-speech acoustic details, and (2) downstream task consistency, which ensures minimal deviation in downstream signal processing tasks when using reconstructed speech instead of the original. Our comprehensive experiments reveal that complex acoustic environments significantly degrade signal reconstruction and downstream task consistency. This work highlights the limitations of current speech codecs and raises a future direction that improves them for greater environmental resilience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。