arXiv:2505.08215cs.AIcs.SD2025-05被引 6

优化语音大模型预测听障者语音可懂度的实用方法

Unveiling the Best Practices for Applying Speech Foundation Models to Speech Intelligibility Prediction for Hearing-Impaired People

  • 选单层编码器比用所有层效果更好
  • 带时序建模的预测头能显著提升准确率
  • 多个大模型集成可进一步提高性能,强者贡献更大

语音基础模型(SFMs)在多种下游任务中表现优异,包括听障人群语音可懂度预测(SIP-HI)。然而,针对SIP-HI优化SFMs的研究仍不充分。本文对5个SFMs进行了全面研究,聚焦编码器层选择、预测头结构和集成配置等关键设计因素。结果表明,与传统全层使用方法相反,仅选用单一编码器层可获得更优效果;时序建模对预测头性能至关重要;多模型集成可提升性能,且个体模型越强,集成收益越大。此外,我们分析了SFMs关键属性与其在SIP-HI任务中表现的关系。本研究为高效适配语音基础模型用于听障人群语音可懂度预测提供了实用指导。

原文摘要 · Abstract (English)

Speech foundation models (SFMs) have demonstrated strong performance across a variety of downstream tasks, including speech intelligibility prediction for hearing-impaired people (SIP-HI). However, optimizing SFMs for SIP-HI has been insufficiently explored. In this paper, we conduct a comprehensive study to identify key design factors affecting SIP-HI performance with 5 SFMs, focusing on encoder layer selection, prediction head architecture, and ensemble configurations. Our findings show that, contrary to traditional use-all-layers methods, selecting a single encoder layer yields better results. Additionally, temporal modeling is crucial for effective prediction heads. We also demonstrate that ensembling multiple SFMs improves performance, with stronger individual models providing greater benefit. Finally, we explore the relationship between key SFM attributes and their impact on SIP-HI performance. Our study offers practical insights into effectively adapting SFMs for speech intelligibility prediction for hearing-impaired populations.

语音可懂度听障辅助大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。