用语义和视觉信息提升低比特率语音编码质量
Enhancing Neural Speech Coding with Semantic and Visual Cues

- 引入语义与视觉分支,通过交叉注意力融合上下文信息
- 在低码率下使语音重建的ViSQOL得分从3.86提升至4.01
- 支持有无辅助信息的两种推理模式,适用场景更灵活
在低比特率下,仅依赖语音特征的神经语音编解码器难以完整保留高质量重建所需信息。本文提出一种语义与视觉增强语音编解码器(SVSC),将语义和视觉线索融入神经语音编码过程。基于主流架构,SVSC引入语义编码-解码分支和图像分析-合成分支,通过交叉注意力机制融合深层语义特征与视觉线索,生成富含上下文和发音信息的高阶辅助表示。针对不同推理场景,设计两种信息注入策略:当辅助线索可用时,采用特征拼接直接融合;否则在训练中通过知识蒸馏将辅助信息迁移至语音编码分支,实现无需额外输入的纯语音推理。实验表明,该方法有效提升了语音重建质量,使ViSQOL得分从3.86提升至4.01。
原文摘要 · Abstract (English)
At low bitrates, neural speech codecs have limited capacity to encode all information needed for high-quality re construction, especially when relying solely on speech-derived representations. To address this limitation, this paper proposes a Semantic- and Visual-enhanced Speech Codec (SVSC), which in corporates semantic and visual cues into the neural speech coding process. Specifically, built upon a mainstream neural speech cod ing architecture, SVSC introduces a semantic encoding-decoding branch and an image analysis-synthesis branch. It fuses deep semantic features with visual cues through a cross-attention mech anism, forming an auxiliary high-level representation enriched with contextual and articulatory information. To handle different inference scenarios, SVSC introduces two information-injection strategies based on the availability of auxiliary semantic and vi sual cues. When such cues are available, the fusion mode directly incorporates the auxiliary representations into the speech coding branch through feature concatenation; otherwise, the distillation mode transfers auxiliary information into the speech coding branch through knowledge distillation during training, enabling speech-only inference without additional inputs. Experimental results validate the effectiveness of incorporating semantic and visual cues, improving the ViSQOL score of reconstructed speech from 3.86 to 4.01.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。