arXiv:2506.09549eess.AScs.SD2025-06中稿 · Interspeech 2025被引 2

用视觉线索提升语音质量评估精度,无需参考音频。

A Study on Speech Assessment with Visual Cues

  • 融合音频频谱与视觉特征,双分支结构协同建模。
  • 在含噪语音上,PESQ和STOI指标分别提升9.61%和11.47%。
  • 适合做无参考语音评估的场景,如视频会议、远程医疗。

当缺乏干净参考信号时,非侵入式语音质量与可懂度评估至关重要。本文提出一种多模态框架,结合音频特征与视觉线索,预测PESQ和STOI得分。采用双分支架构:音频部分通过STFT提取频谱特征,视觉部分通过视觉编码器获取嵌入表示。二者融合后经由CNN-BLSTM加注意力机制处理,并通过多任务学习同时预测两个指标。在LRS3-TED数据集上进行实验,噪声来自DEMAND语料库。结果显示,模型优于仅使用音频的基线。在已见噪声条件下,PESQ的LCC提升9.61%(0.8397→0.9205),STOI提升11.47%(0.7403→0.8253)。结果表明,引入视觉线索能有效提升非侵入式语音评估的准确性。

原文摘要 · Abstract (English)

Non-intrusive assessment of speech quality and intelligibility is essential when clean reference signals are unavailable. In this work, we propose a multimodal framework that integrates audio features and visual cues to predict PESQ and STOI scores. It employs a dual-branch architecture, where spectral features are extracted using STFT, and visual embeddings are obtained via a visual encoder. These features are then fused and processed by a CNN-BLSTM with attention, followed by multi-task learning to simultaneously predict PESQ and STOI. Evaluations on the LRS3-TED dataset, augmented with noise from the DEMAND corpus, show that our model outperforms the audio-only baseline. Under seen noise conditions, it improves LCC by 9.61% (0.8397->0.9205) for PESQ and 11.47% (0.7403->0.8253) for STOI. These results highlight the effectiveness of incorporating visual cues in enhancing the accuracy of non-intrusive speech assessment.

语音评估多模态视觉线索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。