用可控风格生成增强语音,提升说话人聚类对内在变化的鲁棒性。
Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
- 通过可控风格生成模型在语音中引入多样风格,保持说话人身份不变。
- 在模拟情感数据集上错误率降低49%,在AMI数据集上降低35%。
- 适合处理情绪、健康等变化导致的说话人识别偏差问题。
说话人聚类系统常因说话人内部的固有变异性(如情绪、健康或内容变化)而表现不佳,例如声音提高或语速加快时,同一说话人的片段可能被误判为不同个体。为此,我们提出一种可控风格的语音生成模型,在保留目标说话人身份的前提下,对语音进行多风格增强。该系统首先使用传统聚类器得到的说话人分段,对每个分段生成包含音素与风格多样性的增强语音样本,并将原始音频与生成音频的说话人嵌入进行融合,从而提升系统在高内在变异性情况下的分段聚类鲁棒性。我们在模拟情感语音数据集和截断版AMI数据集上验证了该方法,分别实现了49%和35%的错误率下降。
原文摘要 · Abstract (English)
Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different individuals, for example, when one raises their voice or speaks faster during conversation. To address this, we propose a style-controllable speech generation model that augments speech across diverse styles while preserving the target speaker's identity. The proposed system starts with diarized segments from a conventional diarizer. For each diarized segment, it generates augmented speech samples enriched with phonetic and stylistic diversity. And then, speaker embeddings from both the original and generated audio are blended to enhance the system's robustness in grouping segments with high intrinsic intra-speaker variability. We validate our approach on a simulated emotional speech dataset and the truncated AMI dataset, demonstrating significant improvements, with error rate reductions of 49% and 35% on each dataset, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。