arXiv:2506.01483eess.AScs.SD2025-06中稿 · Interspeech 2025被引 3

用说话人相对特征提升文本引导语音分离效果

Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction

  • 基于说话人间相对差异构建连续与离散特征
  • 性别和时间顺序在多语言混响环境下最有效
  • 适配预训练模型,适合语音分离与数据扩展场景

我们提出一种新方法,利用说话人之间的相对线索区分目标说话人并从混合语音中提取其声音。连续线索(如时间顺序、年龄、音高水平)按相对差异分组,离散线索(如语言、性别、情绪)保持类别区分。相比固定语音属性分类,相对线索更具灵活性,更易扩展文本引导语音提取数据集。实验表明,融合所有相对线索的性能优于随机子集,其中性别和时间顺序在多种语言和混响条件下表现最稳健。额外线索如音高水平、音量、距离、发言时长、语言和音域在复杂场景中也显著提升效果。微调预训练WavLM Base+ CNN编码器的表现优于Conv1d基线。

原文摘要 · Abstract (English)

We propose a novel approach that utilizes inter-speaker relative cues to distinguish target speakers and extract their voices from mixtures. Continuous cues (e.g., temporal order, age, pitch level) are grouped by relative differences, while discrete cues (e.g., language, gender, emotion) retain their categorical distinctions. Compared to fixed speech attribute classification, inter-speaker relative cues offer greater flexibility, facilitating much easier expansion of text-guided target speech extraction datasets. Our experiments show that combining all relative cues yields better performance than random subsets, with gender and temporal order being the most robust across languages and reverberant conditions. Additional cues, such as pitch level, loudness, distance, speaking duration, language, and pitch range, also demonstrate notable benefits in complex scenarios. Fine-tuning pre-trained WavLM Base+ CNN encoders improves overall performance over the Conv1d baseline.

语音分离相对特征文本引导多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。