arXiv:2509.03013eess.AScs.SD2025-09中稿 · APSIPA ASC 2025

用不确定性感知的语音嵌入提升智能语音可懂度预测准确率

Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM

  • 用Whisper嵌入的均值、方差和熵构建不确定性感知特征
  • iMTI-Net在多指标上优于原模型,人机可懂度预测更准
  • 适合做语音质量评估与自动语音识别系统优化的研究者

非侵入式语音可懂度预测因说话人差异、噪声环境及主观感知变化而具挑战性。本文提出一种不确定性感知方法,利用Whisper嵌入结合统计特征(均值、标准差、熵),其中熵通过特征维度上的softmax计算,作为不确定性的代理指标,补充均值与标准差所捕捉的全局信息。为建模语音的时序结构,采用标量长短期记忆网络(sLSTM)高效捕获长程依赖。在此基础上,提出iMTI-Net——一种融合卷积神经网络(CNN)与sLSTM的多任务学习框架,联合预测人工可懂度评分与谷歌ASR及Whisper的词错误率(WER)。实验表明,iMTI-Net在多个评估指标上超越原始MTI-Net,验证了不确定性感知特征与CNN-sLSTM架构的有效性。

原文摘要 · Abstract (English)

Non-intrusive speech intelligibility prediction remains challenging due to variability in speakers, noise conditions, and subjective perception. We propose an uncertainty-aware approach that leverages Whisper embeddings in combination with statistical features, specifically the mean, standard deviation, and entropy computed across the embedding dimensions. The entropy, computed via a softmax over the feature dimension, serves as a proxy for uncertainty, complementing global information captured by the mean and standard deviation. To model the sequential structure of speech, we adopt a scalar long short-term memory (sLSTM) network, which efficiently captures long-range dependencies. Building on this foundation, we propose iMTI-Net, an improved multi-target intelligibility prediction network that integrates convolutional neural network (CNN) and sLSTM components within a multitask learning framework. It jointly predicts human intelligibility scores and machine-based word error rates (WER) from Google ASR and Whisper. Experimental results show that iMTI-Net outperforms the original MTI-Net across multiple evaluation metrics, demonstrating the effectiveness of incorporating uncertainty-aware features and the CNN-sLSTM architecture.

语音可懂度Whisper多任务学习sLSTM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。