arXiv:2501.13000eess.AScs.SD2025-01中稿 · ICASSP 2025被引 4

揭秘语音去身份系统如何误删情感信息,揭示核心原因与评估漏洞

Why disentanglement-based speaker anonymization systems fail at preserving emotions?

  • 分析中间表示中缺乏情感信息是主因
  • 生成式训练的说话人嵌入会意外改变情感
  • 提醒用未加权召回率评估情感识别不靠谱

基于解耦的语音去身份技术将语音分解为语义表示,修改说话人嵌入,并通过神经声码器重建波形。当前顶尖系统常导致情感信息丢失。可能原因包括基于GAN的声码器出现模式崩溃、说话人嵌入无意建模或修改情感,或中间表示过度净化。本文对一个顶尖系统进行全面评估,发现主要原因是中间表示缺乏情感相关信息;说话人嵌入若在生成式上下文中学习,影响也很大;声码器分布外性能影响较小。此外,发现合成伪影会提高谱峭度,使情感识别评估偏向将语音判为愤怒。因此,仅报告情感识别的未加权平均召回率是不充分的。

原文摘要 · Abstract (English)

Disentanglement-based speaker anonymization involves decomposing speech into a semantically meaningful representation, altering the speaker embedding, and resynthesizing a waveform using a neural vocoder. State-of-the-art systems of this kind are known to remove emotion information. Possible reasons include mode collapse in GAN-based vocoders, unintended modeling and modification of emotions through speaker embeddings, or excessive sanitization of the intermediate representation. In this paper, we conduct a comprehensive evaluation of a state-of-the-art speaker anonymization system to understand the underlying causes. We conclude that the main reason is the lack of emotion-related information in the intermediate representation. The speaker embeddings also have a high impact, if they are learned in a generative context. The vocoder's out-of-distribution performance has a smaller impact. Additionally, we discovered that synthesis artifacts increase spectral kurtosis, biasing emotion recognition evaluation towards classifying utterances as angry. Therefore, we conclude that reporting unweighted average recall alone for emotion recognition performance is suboptimal.

语音去身份情感保留模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。