arXiv:2503.08798cs.SDcs.LG2025-03中稿 · ICASSP 2025被引 5

仅用对话历史提取目标语音,准确率超90%

Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction

  • 利用对话历史作为隐式线索,无需录音或人脸信息
  • 仅需两轮对话历史,目标语音识别准确率超90%
  • 支持文本与录音双模式推理,适合移动端语音应用

本文提出一种新型目标语音提取方法——上下文语音提取(CSE),仅依赖先前对话的文本内容来定位目标语音流。不同于传统方法需要预录语音、目标说话人面部视频或空间信息等显式线索,本方法仅需少量历史对话文本即可完成定位。该方法在移动消息场景中尤为适用,因语音常伴随文本对话。我们设计了三种CSE模型,在三个数据集上进行评估。实验表明,即使仅依赖对话历史,模型在仅有两轮对话的情况下仍能实现超过90%的正确率。此外,通过在训练中同时使用文本上下文和预录语音作为提示,可提升模型灵活性,使其在推理时可选择任一提示或两者结合以获得更优效果。代码与样例见 https://miraodasilva.github.io/cse-project-page。

原文摘要 · Abstract (English)

In this paper, we investigate a novel approach for Target Speech Extraction (TSE), which relies solely on textual context to extract the target speech. We refer to this task as Contextual Speech Extraction (CSE). Unlike traditional TSE methods that rely on pre-recorded enrollment utterances, video of the target speaker's face, spatial information, or other explicit cues to identify the target stream, our proposed method requires only a few turns of previous dialogue (or monologue) history. This approach is naturally feasible in mobile messaging environments where voice recordings are typically preceded by textual dialogue that can be leveraged implicitly. We present three CSE models and analyze their performances on three datasets. Through our experiments, we demonstrate that even when the model relies purely on dialogue history, it can achieve over 90 % accuracy in identifying the correct target stream with only two previous dialogue turns. Furthermore, we show that by leveraging both textual context and enrollment utterances as cues during training, we further enhance our model's flexibility and effectiveness, allowing us to use either cue during inference, or combine both for improved performance. Samples and code available on https://miraodasilva.github.io/cse-project-page .

语音提取上下文理解多模态移动端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。