arXiv:2606.25959eess.AScs.AI2026-06中稿 · Interspeech 2026

端到端联合优化语音增强与音量控制,提升会议场景音质和识别率。

SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios

论文配图:SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios
图 1 · 摘自论文原文
  • 端到端联合优化语音增强与自动增益控制,利用任务间协同效应。
  • 在会议场景中实现目标音量,同时提升语音质量和语音识别准确率。
  • 适用于语音处理、智能会议系统等需要稳定音量的场景。

传统音频处理流程将语音增强(SE)与自动增益控制(AGC)视为独立模块,常导致整体性能受限。例如,在增强前进行AGC可能放大背景噪声,而优先进行SE又易过度抑制低音量语音。为此,本文提出SE-AGCNet,一种针对会议场景中显著音量差异的端到端联合优化框架。该框架利用两任务间的协同效应:SE保留低音量语音,使AGC能更有效调节音量。此外,我们设计了专用数据生成管道SE-AGC-DataGen,并引入标准化响度评估指标:总响度(LUFS)、短时响度(St LUFS)和响度动态范围(LRA)。实验表明,SE-AGCNet在保持目标响度的同时,优于现有基线方法,在语音质量与语音识别准确率方面均有提升。

原文摘要 · Abstract (English)

Conventional audio pipelines typically treat speech enhancement (SE) and automatic gain control (AGC) as discrete modules, which often limits overall performance. For instance, applying AGC before SE may inadvertently amplify background noise, while prioritizing SE tends to over-suppress low-volume speech. To address these limitations, we propose SE-AGCNet, an end-to-end framework that jointly optimizes SE and AGC. Tailored for meeting scenarios with significant volume variations, SE-AGCNet leverages the synergy between the two tasks: SE preserves quiet speech, thereby facilitating effective volume adjustment by the AGC component. Furthermore, we propose a specialized data simulation pipeline, SE-AGC-DataGen, and incorporate standardized loudness evaluation metrics: integrated loudness (LUFS), short-term loudness (St LUFS), and LRA. Experiments show that SE-AGCNet consistently achieves target loudness while improving speech quality and ASR accuracy over competitive baselines.

语音增强自动增益会议系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。