arXiv:2501.09169eess.AScs.SD2025-01中稿 · ICASSP 2025被引 10

用文字描述说话风格,实现无语音样本的精准语音提取

Beyond Speaker Identity: Text Guided Target Speech Extraction

  • 结合语音与文本双重线索,用自然语言描述风格辅助分离目标语音
  • 在无录音样本场景下,语音分离准确率提升23.6%(相对提升)
  • 适合语音识别、智能客服等缺乏说话人音频的实用场景

传统目标语音提取(TSE)依赖说话人录音、人脸图像等显式身份信息,但这些信息常不可得。本文提出文本引导的TSE模型StyleTSE,利用自然语言对说话风格的描述,结合音频线索从混叠语音中提取目标语音。模型融合基于SepFormer的语音分离网络与双模态线索网络,可灵活处理音频与文本输入。为训练与评估,我们构建了新数据集TextrolMix,包含语音混合对与自然语言描述。实验表明,该方法不仅能依据说话人身份,还能根据说话风格有效分离语音,在传统音频线索缺失时显著提升性能。演示地址:https://mingyue66.github.io/TextrolMix/demo/

原文摘要 · Abstract (English)

Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker's identity like enrollment audio, face images, or videos, which may not always be available. In this paper, we propose a text-guided TSE model StyleTSE that uses natural language descriptions of speaking style in addition to the audio clue to extract the desired speech from a given mixture. Our model integrates a speech separation network adapted from SepFormer with a bi-modality clue network that flexibly processes both audio and text clues. To train and evaluate our model, we introduce a new dataset TextrolMix with speech mixtures and natural language descriptions. Experimental results demonstrate that our method effectively separates speech based not only on who is speaking, but also on how they are speaking, enhancing TSE in scenarios where traditional audio clues are absent. Demos are at: https://mingyue66.github.io/TextrolMix/demo/

语音分离文本引导多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。