arXiv:2506.09792cs.SDcs.LG2025-06中稿 · Interspeech 2025被引 4

用语言模型提升音视频语音分离效果,不增加推理成本

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

  • 引入预训练语言模型作为语言约束辅助语音分离
  • 在多语言和视觉线索弱的场景下均显著提升音质与可懂度
  • 无需额外计算开销,适合实际部署

音视频目标说话人分离(AV-TSE)模型主要依赖目标说话人的视觉线索来分离其语音。人类在听觉感知中会借助语法、语义等语言知识。受此启发,我们探索将预训练语音-语言模型(PSLMs)和预训练语言模型(PLMs)作为辅助知识源用于AV-TSE。本文提出将PSLM或PLM的语言约束融入AV-TSE模型,作为额外监督信号。该方法在推理阶段不引入额外计算开销,能持续提升语音质量和可懂度。此外,在多语言设置和视觉线索受损场景下也表现出稳健的性能提升。

原文摘要 · Abstract (English)

Audio-visual target speaker extraction (AV-TSE) models primarily rely on target visual cues to isolate the target speaker's voice from others. We know that humans leverage linguistic knowledge, such as syntax and semantics, to support speech perception. Inspired by this, we explore the potential of pre-trained speech-language models (PSLMs) and pre-trained language models (PLMs) as auxiliary knowledge sources for AV-TSE. In this study, we propose incorporating the linguistic constraints from PSLMs or PLMs for the AV-TSE model as additional supervision signals. Without introducing any extra computational cost during inference, the proposed approach consistently improves speech quality and intelligibility. Furthermore, we evaluate our method in multi-language settings and visual cue-impaired scenarios and show robust performance gains.

语音分离语言模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。