arXiv:2605.11098cs.SD2026-05ACL被引 1

让语音编码器在压缩时保留情感信息,提升表达自然度。

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling

论文配图:AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
图 1 · 摘自论文原文
  • 用情感引导的潜空间调制,保留关键情绪特征。
  • 在重建和生成任务中,情感一致性提升且内容准确无损。
  • 适合需要高情感表达的语音合成与交互系统。

神经语音编码器为语音语言模型提供离散表示,但量化过程常导致情感线索丢失。现有编码器主要优化声学重建,未能充分建模表示层面的情感表现力。本文提出一种情感引导的神经语音编码器,通过情感-语义联合调制、关系保持的情感-语义蒸馏以及情感加权的语义对齐,在压缩下显式保留情感显著特征,同时维持语义一致性和韵律自然性。在语音重建、情感识别及下游文本到语音生成任务上的广泛评估表明,该方法在不牺牲内容准确性的前提下,显著提升了情感一致性和感知质量。

原文摘要 · Abstract (English)

Neural speech codecs provide discrete representations for speech language models, but emotional cues are often degraded during quantization. Existing codecs mainly optimize acoustic reconstruction, leaving emotion expressiveness insufficiently modeled at the representation level. We propose an emotion-guided neural speech codec that explicitly preserves emotional information while maintaining semantic fidelity and prosodic naturalness. Our framework combines emotion-semantic guided latent modulation, relation-preserving emotional-semantic distillation, and emotion-weighted semantic alignment to retain emotionally salient cues under compression. Extensive evaluations across speech reconstruction, emotion recognition, and downstream text-to-speech generation demonstrate improved emotion consistency and perceptual quality without sacrificing content accuracy.

语音编码情感建模语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。