分离情绪干扰,提升语音聚类准确率
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
- 在变分自编码器中设计解耦框架,分离情绪与说话人特征
- 在多个数据集上使聚类准确率显著提升,最高达12.3%
- 适合需要高鲁棒性语音聚类的场景,如会议记录自动标注
说话人聚类旨在从一组音频中识别出唯一的说话人(每个音频仅属于一个说话人),而无需事先知道总共有多少说话人,是说话人区分任务中的关键步骤。近年来,通用深度说话人嵌入模型被广泛用于捕捉说话人特征。然而,带有情感表达的语音会带来显著挑战,常影响说话人嵌入质量,导致聚类性能下降。为此,我们提出DTG-VAE,一种新型解耦方法,在变分自编码器(VAE)框架内增强聚类能力。本研究揭示了情绪状态与深度说话人嵌入效果之间的直接关联。实验表明,DTG-VAE能提取更鲁棒的说话人嵌入,显著提升说话人聚类性能。
原文摘要 · Abstract (English)
Speaker clustering is the task of identifying the unique speakers in a set of audio recordings (each belonging to exactly one speaker) without knowing who and how many speakers are present in the entire data, which is essential for speaker diarization processes. Recently, off-the-shelf deep speaker embedding models have been leveraged to capture speaker characteristics. However, speeches containing emotional expressions pose significant challenges, often affecting the accuracy of speaker embeddings and leading to a decline in speaker clustering performance. To tackle this problem, we propose DTG-VAE, a novel disentanglement method that enhances clustering within a Variational Autoencoder (VAE) framework. This study reveals a direct link between emotional states and the effectiveness of deep speaker embeddings. As demonstrated in our experiments, DTG-VAE extracts more robust speaker embeddings and significantly enhances speaker clustering performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。