arXiv:2608.23759eess.AScs.SD2026-08

评测真实场景下视听语音增强效果,发布基准模型与评估工具。

The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge

论文配图:The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
图 1 · 摘自论文原文
  • 设计真实混叠与视觉失效的双轨挑战任务
  • 基准模型在合成混音上达SI-SDR -2.851dB、STOI 0.470
  • 适合语音增强与多模态信号处理研究者参考

视听语音增强(AVSE)利用目标说话人的视觉信息从噪声或重叠语音中恢复其语音。现有广泛使用的评估协议通常基于独立录制的音频混合,并假设视频质量可靠,导致对真实重叠与视觉失效情况下的性能刻画不足。本挑战评估两个相关设置:第一轨包含两种场景——双人同时说话的真实混合信号(无干净参考),以及手动混合的合成混音(有干净参考);第二轨使用相同音频但配以退化的目标视频,并新增3米远场录音。开发集与测试集说话人互不重叠。评估指标包括波形保真度、学习质量估计、语音识别准确率和说话人识别。在开发集的合成混音任务中,基线模型在第一轨达到SI-SDR -4.069 dB、STOI 0.388,第二轨为SI-SDR -2.851 dB、STOI 0.470。我们公开了AV-ConvTasNet检查点、离线评估器及开发集与测试集的官方基线结果。

原文摘要 · Abstract (English)

Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols construct mixed signals from separately recorded audio sources and assume reliable video, leaving their performance under natural overlap and visual failure insufficiently characterized. The Real-World AVSE Challenge evaluates two related settings. Track~1 comprises two scenarios: real-world mixtures recorded with two speakers speaking simultaneously, without a corresponding clean reference signal, and synthetic remixes obtained by manually mixing the separately recorded speech of two speakers, with a clean reference signal available; Track~2 reuses audio but pairs it with a degraded target video and contains additional 3-m far-field recordings. The speakers in the development and test sets are disjoint. Evaluation metrics include clean-waveform fidelity, learned quality estimates, transcription accuracy, and speaker identification. In the remix task on the development set, the baseline model achieved an SI-SDR of $-4.069$~dB and an STOI of $0.388$ on Track~1, and an SI-SDR of $-2.851$~dB and an STOI of $0.470$ on Track~2. We release the AV-ConvTasNet checkpoints, the offline evaluator, and the official baseline results on the development and test sets.

语音增强视听融合真实场景挑战赛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。