用已有听音成绩预测个人语音可懂度,比传统测听更准。
No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction
- 基于多组音频与评分构建听者识别能力的高维表征
- 少量支持样本下预测精度超越传统测听数据
- 适合语音可懂度个性化评估的研究与应用
个性化语音可懂度预测面临挑战。以往方法主要依赖听觉图谱,但其仅反映纯音听阈,精度有限。本文提出新方法:利用个体已有的语音可懂度数据,预测其对新音频的表现。我们设计支持样本驱动的可懂度预测网络(SSIPNet),通过语音基础模型从多个(音频, 分数)对中构建听者语音识别能力的高维表征,实现对未见音频的精准预测。在Clarity Prediction Challenge数据集上的实验表明,即使仅有少量支持样本,该方法仍优于基于听觉图谱的预测。本工作为个性化语音可懂度预测提供了新范式。
原文摘要 · Abstract (English)
Personalized speech intelligibility prediction is challenging. Previous approaches have mainly relied on audiograms, which are inherently limited in accuracy as they only capture a listener's hearing threshold for pure tones. Rather than incorporating additional listener features, we propose a novel approach that leverages an individual's existing intelligibility data to predict their performance on new audio. We introduce the Support Sample-Based Intelligibility Prediction Network (SSIPNet), a deep learning model that leverages speech foundation models to build a high-dimensional representation of a listener's speech recognition ability from multiple support (audio, score) pairs, enabling accurate predictions for unseen audio. Results on the Clarity Prediction Challenge dataset show that, even with a small number of support (audio, score) pairs, our method outperforms audiogram-based predictions. Our work presents a new paradigm for personalized speech intelligibility prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。