arXiv:2410.11243cs.SDcs.CL2024-10中稿 · IEEE SLT 2024被引 2

对比不同说话人嵌入,发现理想嵌入能显著提升目标说话人语音处理效果

Investigation of Speaker Representation for Target-Speaker Speech Processing

  • 用预训练编码器和理想嵌入对比,统一评估三类任务
  • 一热向量嵌入优于录音采集的嵌入,性能更优
  • 最优嵌入依赖输入混合信号,适合语音分离与识别场景

目标说话人语音处理(TS)任务如目标说话人自动语音识别(TS-ASR)、目标语音提取(TSE)和个人语音活动检测(p-VAD),旨在从混杂语音中提取特定说话人的信息。尽管现有研究多聚焦于各任务的训练策略或系统架构,但对用于编码目标说话人线索的辅助网络尚未在跨任务统一评估中深入探讨。本文旨在回答一个基础问题:何种说话人嵌入最适合TS任务?针对TS-ASR、TSE和p-VAD,我们比较了基于预录制目标说话人语音的预训练说话人编码器(自监督或说话人识别模型)生成的嵌入,与直接由目标说话人身份生成的一热向量(one-hot vector)形式的理想嵌入。为进一步理解理想嵌入特性,我们采用基于梯度的方法优化其表示以提升任务性能。分析表明,说话人验证性能与TS任务表现相关性较弱,一热向量优于基于录音的嵌入,且最优嵌入取决于输入语音混合情况。

原文摘要 · Abstract (English)

Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-VAD), are important for extracting information about a desired speaker's speech even when it is corrupted by interfering speakers. While most studies have focused on training schemes or system architectures for each specific task, the auxiliary network for embedding target-speaker cues has not been investigated comprehensively in a unified cross-task evaluation. Therefore, this paper aims to address a fundamental question: what is the preferred speaker embedding for TS tasks? To this end, for the TS-ASR, TSE, and p-VAD tasks, we compare pre-trained speaker encoders (i.e., self-supervised or speaker recognition models) that compute speaker embeddings from pre-recorded enrollment speech of the target speaker with ideal speaker embeddings derived directly from the target speaker's identity in the form of a one-hot vector. To further understand the properties of ideal speaker embedding, we optimize it using a gradient-based approach to improve performance on the TS task. Our analysis reveals that speaker verification performance is somewhat unrelated to TS task performances, the one-hot vector outperforms enrollment-based ones, and the optimal embedding depends on the input mixture.

说话人嵌入语音分离语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。