arXiv:2411.03109cs.SDcs.MM2024-11

用零散文字提示提取会议演讲中的目标人声,效果显著。

pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues

  • 用不齐步的文本内容生成语义提示,指导语音分离
  • 在无对齐文本下实现12.16 dB的SI-SDRi性能
  • 适合会议、海报展示等难获取音频线索的场景

目标说话人分离(TSE)旨在从混叠音频中提取目标说话人的清晰语音,消除无关噪声和干扰语音。以往方法依赖预录语音、视觉信息或空间信息等强辅助线索,但在许多实际场景中难以实时获取。本文提出一种新方法,利用有限且未对齐的文本内容(如演示文稿要点)作为语义提示来引导TSE。为此,设计了文本提示提取网络(TPE),将音频特征与基于内容的语义线索融合,生成时频掩码以过滤噪声。实验表明,仅使用有限且未对齐的文本提示,即可实现优异的语音分离效果:在测试集上达到SI-SDRi 12.16 dB、SDRi 12.66 dB、PESQi 0.830 和 STOIi 0.150。

原文摘要 · Abstract (English)

Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has explored various auxiliary cues including pre-recorded speech, visual information, and spatial information, the acquisition and selection of such strong cues are infeasible in many practical scenarios. Differently, in this paper, we condition the TSE algorithm on semantic cues extracted from limited and unaligned text contents, such as condensed points from a presentation slide. This method is particularly useful in scenarios like meetings, poster sessions, or lecture presentations, where acquiring other cues in real time may be challenging. To this end, we design two different networks. Specifically, our proposed Text Prompt Extractor Network (TPE) fuses audio features with content-based semantic cues to facilitate time-frequency mask generation to filter out extraneous noise. The experimental results show the efficacy in accurately extracting the target speaker's speech by utilizing semantic cues derived from limited and unaligned text, resulting in SI-SDRi of 12.16 dB, SDRi of 12.66 dB, PESQi of 0.830 and STOIi of 0.150.

语音分离文本提示会议记录

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。