用帧对齐融合双模型,提升助听器语音可懂度预测精度
Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech
- 在粗粒度时间线上对齐并融合Canary与WavLM特征
- 达到评估RMSE 24.96、相关系数0.796的性能
- 适合语音可懂度评估与助听器算法研究者
非侵入式可懂度预测旨在不依赖干净参考信号的情况下估计听力障碍者对助听器处理语音的理解程度。本文在第三届清晰度预测挑战赛中,采用两个冻结的语音编码器Canary和WavLM进行研究。核心问题是:是否应融合互补的预训练表示,以及融合应在何处发生。我们对比了单骨干基线、统一评分平均、池化后融合、交叉注意力、帧对齐融合及反向对齐等方法,在共享左右保持的双耳框架下进行评估。最佳模型通过可学习的步进卷积对WavLM进行时间准备,并在较粗的Canary时间线上与Canary特征融合后再池化,最终在测试集上取得RMSE 24.96±0.06、相关系数0.796±0.001的成绩。严重程度、增强系统、层窗和时间偏移分析表明,池化前的粗粒度局部时间对应是一种有效的归纳偏置。
原文摘要 · Abstract (English)
Non-intrusive intelligibility prediction estimates how well hearing-impaired listeners understand hearing-aid-processed speech without a clean reference. We study this task in the 3rd Clarity Prediction Challenge using two frozen speech encoders, Canary and WavLM. The central question is not only whether complementary pretrained representations should be combined, but where their interaction should occur. We compare single-backbone baselines, uniform score averaging, pool-late fusion, cross-attention, frame-aligned fusion, and reverse alignment under a shared left/right-preserving binaural framework. Among the compared systems, the best model temporally prepares WavLM with a learnable strided convolution and fuses it with Canary on the coarser Canary timeline before pooling, reaching Eval RMSE 24.96$\pm$0.06 and Eval Corr 0.796$\pm$0.001. Severity, enhancement-system, layer-window, and temporal-shift analyses indicate that coarse local temporal correspondence before pooling is a useful inductive bias for this task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。