用相对线索提升语音分离精度,效果优于传统方法。
Inter-Speaker Relative Cues for Two-Stage Text-Guided Target Speech Extraction
- 分两阶段:先生成候选语音,再用文本匹配选目标说话人
- 相对线索使分类准确率和语音提取性能均更高
- 适合需要精准区分说话人场景的语音处理应用
本文研究了在文本引导的目标语音提取(TSE)中使用相对线索的有效性。从人类感知和标签量化角度出发,理论证明相对线索能保留绝对类别表示中丢失的细粒度差异,尤其适用于连续属性。基于此,提出两阶段TSE框架:首先由语音分离模型生成候选语音源,随后通过文本引导分类器依据嵌入相似性筛选目标说话人。训练两个独立分类模型,对比相对线索与独立线索在连续属性下的表现,涵盖分类准确率和TSE性能。实验表明:(i) 相对线索在整体分类准确率和TSE性能上均优于独立线索;(ii) 所提两阶段框架在信号级和客观感知指标上显著超越单阶段文本条件提取方法;(iii) 多种相对线索(包括语言、音量、距离、时间顺序、发言时长、随机线索及所有线索)甚至超过基于注册音频的TSE系统表现。进一步分析揭示不同线索类型间存在显著判别力差异,为相对线索在TSE中的有效性提供洞见。
原文摘要 · Abstract (English)
This paper investigates the use of relative cues for text-based target speech extraction (TSE). We first provide a theoretical justification for relative cues from the perspectives of human perception and label quantization, showing that relative cues preserve fine-grained distinctions that are often lost in absolute categorical representations for continuous-valued attributes. Building on this analysis, we propose a two-stage TSE framework in which a speech separation model first generates candidate sources, followed by a text-guided classifier that selects the target speaker based on embedding similarity. Within this framework, we train two separate classification models to evaluate the advantages of relative cues over independent cues in case of continuous-valued attributes, considering both classification accuracy and TSE performance. Experimental results demonstrate that (i) relative cues achieve higher overall classification accuracy and improved TSE performance compared with independent cues; (ii) the proposed two-stage framework substantially outperforms single-stage text-conditioned extraction methods on both signal-level and objective perceptual metrics; and (iii) several relative cues, including language, loudness, distance, temporal order, speaking duration, random cues, and all cues, can even surpass the performance of an enrollment-audio-based TSE system. Further analysis reveals notable differences in discriminative power across cue types, providing insights into the effectiveness of different relative cues for TSE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。