提出AffectCodec,让语音编码在低比特下也能保留情感信息。
AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ

- 用分块对角结构的量化方式,显式保护情感特征空间。
- 在低比特率下情感保留能力显著提升,语音质量仍优秀。
- 适合需要保留情感信息的语音生成、对话系统等应用。
神经语音编解码器已成为原始音频与语音语言模型之间的离散接口,但其主要优化目标仍是声学重建保真度,导致情感相关线索在量化过程中容易丢失,限制了下游模型的情感表达能力。我们发现这一退化源于两个机制:有限比特率下的重建驱动比特分配,以及基于拼接的编解码器中跨流泄漏,使声学梯度覆盖本应保留情感的维度。为此,我们提出AffectCodec,基于分块对角残差有限标量量化(BD-RFSQ),通过在情感与声学子空间上施加分块对角输入输出投影,将比特分配从隐式、损失驱动转变为显式、结构保证,同时保持对下游语音语言模型友好的扁平令牌接口。AffectCodec进一步结合多粒度情感条件建模与多速率训练,在低比特率下实现鲁棒的情感保留。多个情感语音基准测试结果表明,AffectCodec在低比特率下显著提升情感保留性能,同时维持具有竞争力的声学质量和可懂度。结果表明,结构化保护的量化是保留情感相关信息的有效原则,可能为属性感知的神经语音压缩提供通用路径。
原文摘要 · Abstract (English)
Neural speech codecs have become the discrete interface between raw audio and speech language models, yet they remain optimized primarily for acoustic reconstruction fidelity, which leaves emotion-relevant cues vulnerable to being discarded during quantization, limiting the affective capacity of downstream models. We trace this degradation to two mechanisms: reconstruction-driven bit allocation under limited bitrate and cross-stream leakage in concatenation-based codecs, where acoustic gradients can overwrite nominally emotion-reserved dimensions. We propose AffectCodec, an emotion-preserving neural speech codec built on Block-Diagonal Residual Finite Scalar Quantization (BD-RFSQ). By imposing block-diagonal input and output projections over emotion and acoustic subspaces, BD-RFSQ transforms bit allocation from implicit and loss-driven to explicit and structurally guaranteed, while still preserving a flat token interface for downstream speech language models. AffectCodec further combines this structurally constrained quantizer with multi-granularity emotion conditioning and multi-rate training, enabling robust affect preservation at low bitrates. Experiments across multiple emotional speech benchmarks show that AffectCodec substantially improves emotion preservation, especially in the low-bitrate regime, while maintaining competitive acoustic quality and intelligibility. These results suggest that structurally protected quantization is an effective principle for preserving emotion-relevant information and may provide a general route toward attribute-aware neural speech compression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。