提出情感感知量化方法,提升语音离散表示中的情感保留能力。
Emotion-Aware Quantization for Discrete Speech Representations: An Analysis of Emotion Preservation
- 用情感特异性码本实现情感感知量化,改进情绪信息保留。
- 在低比特率下,情感识别准确率显著提升,最高达12.3%。
- 适合需要高保真情感表达的语音合成与情感识别任务。
现代语音系统越来越多地使用离散化的自监督语音表示进行压缩和与基于标记的模型集成,但其对情感信息的影响尚不明确。我们从表示层面和任务层面分析了残差向量量化(RVQ)如何重塑离散语音表示中的情感信息。分析表明,激进压缩会不成比例地破坏情感,不同情感类别和模型架构间存在不均衡损失。为此,我们引入情感感知量化,采用情感特异性和情感偏向的码本,改善硬/软情感感知的保留。进一步提出Emo-Q,一种轻量级路由量化方法,可选择情感专业化码本,在更低比特率下提升情感识别性能。结果强调了情感感知离散化在鲁棒情感语音处理中的重要性。
原文摘要 · Abstract (English)
Modern speech systems increasingly use discretized self-supervised speech representations for compression and integration with token-based models, yet their impact on emotional information remains unclear. We study how residual vector quantization (RVQ) reshapes emotional information in discrete speech representations from both representation- and task-level perspectives. Our analysis shows that aggressive compression disproportionately degrades emotion, with uneven loss across emotion classes and model architectures. To address this, we introduce emotion-aware quantization using emotion-specific and emotion-biased codebooks, improving the preservation of both hard and soft emotion perception. We further propose Emo-Q, a lightweight routed quantization method that selects emotion-specialized codebooks, improving emotion recognition performance at lower bitrates. These results highlight the importance of emotion-aware discretization for robust affective speech processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。