arXiv:2509.17270eess.AScs.SD2025-09

用参考信号增强语音基础模型,提升语音可懂度预测精度

Reference-aware SFM layers for intrusive intelligibility prediction

  • 将参考信号与多层语音基础模型结合,改进侵入式预测
  • 在开发集和测试集上分别达到22.36和24.98的RMSE
  • 为基于语音基础模型的侵入式预测提供实用方案

利用显式参考信号的侵入式语音可懂度预测系统已广泛使用,但其性能尚未持续超越非侵入式系统。我们认为主要原因是语音基础模型(SFM)未被充分挖掘。本文通过引入参考条件与多层SFM表示,重新设计侵入式预测方法。最终系统在开发集上取得22.36的RMSE,评估集上为24.98,在CPC3榜单中排名第一。这些结果为构建基于SFM的侵入式可懂度预测器提供了实践指导。

原文摘要 · Abstract (English)

Intrusive speech-intelligibility predictors that exploit explicit reference signals are now widespread, yet they have not consistently surpassed non-intrusive systems. We argue that a primary cause is the limited exploitation of speech foundation models (SFMs). This work revisits intrusive prediction by combining reference conditioning with multi-layer SFM representations. Our final system achieves RMSE 22.36 on the development set and 24.98 on the evaluation set, ranking 1st on CPC3. These findings provide practical guidance for constructing SFM-based intrusive intelligibility predictors.

语音可懂度参考信号基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。