arXiv:2510.26825cs.SDcs.CV2025-10被引 1

提出分离先于去混响的音视频语音增强新方法,提升复杂环境下的语音质量。

Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling

  • 采用先分离后去混响的流水线设计,提升复杂场景处理能力
  • 在AVSEC-4挑战赛中三项客观指标领先,主观听感排名第一
  • 适用于真实世界多模态复杂声学环境,适合语音增强研究者参考

音视频语音增强(AVSE)利用视觉信息辅助从混合音频中提取目标说话人语音。在真实场景中,常存在复杂的声学环境,伴随多种干扰声音和混响,现有方法难以应对,导致语音感知质量差。本文提出一种高效AVSE系统,可在复杂声学环境中表现优异。具体地,设计了“分离先于去混响”的流水线,可扩展至其他AVSE网络。本方法在第四届COGMHEAR音视频语音增强挑战赛(AVSEC-4)中验证:在竞赛排行榜上三项客观指标均表现优秀,最终在人工主观听感测试中获得第一名。

原文摘要 · Abstract (English)

Audio-visual speech enhancement (AVSE) is a task that uses visual auxiliary information to extract a target speaker's speech from mixed audio. In real-world scenarios, there often exist complex acoustic environments, accompanied by various interfering sounds and reverberation. Most previous methods struggle to cope with such complex conditions, resulting in poor perceptual quality of the extracted speech. In this paper, we propose an effective AVSE system that performs well in complex acoustic environments. Specifically, we design a "separation before dereverberation" pipeline that can be extended to other AVSE networks. The 4th COGMHEAR Audio-Visual Speech Enhancement Challenge (AVSEC) aims to explore new approaches to speech processing in multimodal complex environments. We validated the performance of our system in AVSEC-4: we achieved excellent results in the three objective metrics on the competition leaderboard, and ultimately secured first place in the human subjective listening test.

音视频增强语音分离去混响多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。