构建大规模听觉助听器语音理解易读性数据集与预测模型。
A Large-Scale Database and Predictive Model of Listener-Rated Ease of Speech Understanding in Commercial Hearing Aids
- 用Whisper编码器提取音轨差异特征,训练小模型预测用户感知易读性。
- 新模型在嘈杂场景下相关性达0.89,优于传统HASPIv2的0.75。
- 适合助听器研发、用户体验评估及临床选配参考。
HearAdvisor旨在为听觉助听器消费者提供反映真实聆听体验的音频性能指标与录音。针对语音相关指标,此前主要使用已验证于模拟失真的HASPIv2,但其与真实商业助听器用户感知易读性的关联尚不明确。本文引入一个大规模感知数据集及学习型度量模型,用于预测用户对语音理解易读性的评分。网站访问者(自述听力损失)完成盲测式MUSHRA风格听觉测试,对商业助听器录音在五级“语音理解易读性”量表上打分。数据集包含151,608条评分,经质量筛选后保留104,298条,涵盖83款商业产品在72种真实声学场景中的10,394组双耳声学假人录音。为预测评分,将助听后音频与匹配的纯净语音参考输入冻结的Whisper编码器,提取其内部表示差值,再训练小型MLP头。在训练集外设备上,该模型在场景层面显著优于HASPIv2(整体r=0.92 vs. 0.83;嘈杂场景0.89 vs. 0.75;安静场景0.79 vs. 0.58)。在嘈杂场景中,模型表现达到用户评分的分半信度;在安静场景中接近该上限。模型对增益与信噪比的可控调节也做出合理响应。该数据集与模型为真实商业助听器录音的用户感知易读性提供了全新预测方式。
原文摘要 · Abstract (English)
HearAdvisor aims to provide hearing-aid consumers with audio-performance metrics and recordings that reflect real listening experience. For speech-related metrics, HearAdvisor has historically used HASPIv2, a metric designed to predict objective intelligibility and validated primarily under simulated distortions. Its relationship to consumer-rated ease of understanding for commercial hearing aids is uncertain. Here we introduce a large-scale perceptual dataset and learned metric for listener-rated perceived benefit for speech understanding. Website visitors with self-reported hearing loss completed a blind, MUSHRA-inspired listening test in which they rated recordings of commercial hearing aids on a five-point "Ease of Understanding" scale. The dataset contains 151,608 ratings, 104,298 after quality screening, spanning 10,394 binaural acoustic-manikin recordings from 83 commercial products across 72 realistic acoustic scenes. To predict these ratings, we pass aided audio and a matched clean-speech reference through a frozen Whisper encoder, subtract their internal representations, and train a small MLP head on the resulting difference embedding. On devices held out of training, the learned metric substantially outperforms HASPIv2 at the scene level (overall r = 0.92 vs. 0.83; loud = 0.89 vs. 0.75; quiet = 0.79 vs. 0.58). In loud scenes, performance reaches the split-half reliability of the listener ratings; in quiet scenes, it approaches that ceiling. The model also responds sensibly to controlled gain and SNR manipulations. Together, the dataset and model provide a new way to predict listener-rated ease of speech understanding for real commercial hearing-aid recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。