arXiv:2505.23379eess.AScs.SD2025-05中稿 · interspeech2025

用嘴型信息提升语音编码质量,不增加码率

Vision-Integrated High-Quality Neural Speech Coding

  • 融合唇部图像特征,增强语音编码
  • 在相同码率下,语音质量显著提升
  • 适合语音降噪、远程通信等场景

本文提出一种新型视觉融合神经语音编解码器(VNSC),通过利用视觉模态信息提升语音编码质量。VNSC包含图像分析-合成模块,从唇部图像中提取视觉特征;特征融合模块则实现视觉与语音编码模块间的交互,将视觉信息传递至语音编码过程。根据推理阶段是否具备视觉信息,该模块采用显式融合或隐式蒸馏策略整合视觉特征。实验结果表明,融入视觉信息可有效提升解码语音质量,并增强神经语音编解码器的抗噪能力,且无需增加比特率。

原文摘要 · Abstract (English)

This paper proposes a novel vision-integrated neural speech codec (VNSC), which aims to enhance speech coding quality by leveraging visual modality information. In VNSC, the image analysis-synthesis module extracts visual features from lip images, while the feature fusion module facilitates interaction between the image analysis-synthesis module and the speech coding module, transmitting visual information to assist the speech coding process. Depending on whether visual information is available during the inference stage, the feature fusion module integrates visual features into the speech coding module using either explicit integration or implicit distillation strategies. Experimental results confirm that integrating visual information effectively improves the quality of the decoded speech and enhances the noise robustness of the neural speech codec, without increasing the bitrate.

语音编码多模态视觉融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。