用视线方向提升多人对话中语音增强的准确率。
Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework
- 利用视线方向辅助选择目标说话人,解决多说话人环境中的混淆问题。
- 在多个指标上显著优于无视线引导的基线方法,性能提升超20%。
- 适合需要精准语音分离的智能会议、助听设备等实际场景使用。
本文提出一种基于视线引导的视听语音增强框架(GG-AVSE),以应对鸡尾酒会问题。传统视听语音增强在多说话人环境中难以确定听众关注的目标说话人。GG-AVSE通过引入视线方向作为监督信号,实现目标说话人选择。具体地,提出GG-VM模块,结合YOLO5Face检测器与视线信息提取目标说话人面部特征,并通过零样本融合与部分视觉微调两种策略,与预训练的AVSEMamba模型集成。为评估,构建了新数据集AVSEC2-Gaze。实验表明,相较于无视线引导基线,GG-AVSE在PESQ上提升10.08%(2.370→2.609),STOI提升5.18%(0.8802→0.9258),SI-SDR提升23.69%(9.16→11.33)。结果证实视线能有效缓解目标说话人歧义,且框架具备良好的实际应用扩展性。
原文摘要 · Abstract (English)
This paper presents a Gaze-Guided Audio-Visual Speech Enhancement (GG-AVSE) framework to address the cocktail party problem. A major challenge in conventional AVSE is identifying the listener's intended speaker in multi-talker environments. GG-AVSE addresses this issue by exploiting gaze direction as a supervisory cue for target-speaker selection. Specifically, we propose the GG-VM module, which combines gaze signals with a YOLO5Face detector to extract the target speaker's facial features and integrates them with the pretrained AVSEMamba model through two strategies: zero-shot merging and partial visual fine-tuning. For evaluation, we introduce the AVSEC2-Gaze dataset. Experimental results show that GG-AVSE achieves substantial performance gains over gaze-free baselines: a 10.08% improvement in PESQ (2.370 to 2.609), a 5.18% improvement in STOI (0.8802 to 0.9258), and a 23.69% improvement in SI-SDR (9.16 to 11.33). These results confirm that gaze provides an effective cue for resolving target-speaker ambiguity and highlight the scalability of GG-AVSE for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。