用Whisper提取特征,预测歌词可懂度,效果显著优于基线。
LIWhiz: A Non-Intrusive Lyric Intelligibility Prediction System for the Cadenza Challenge
- 基于Whisper提取声学特征,搭配可训练后端预测分数
- 在CLIP数据集上RMSE为27.07%,相对基线降低22.4%
- 适合语音可懂度评估与音乐语音处理研究者
我们提出LIWhiz,一个提交至ICASSP 2026 Cadenza挑战的非侵入式歌词可懂度预测系统。LIWhiz利用Whisper进行鲁棒特征提取,并采用可训练后端实现评分预测。在Cadenza歌词可懂度预测(CLIP)评测集上,其均方根误差(RMSE)达到27.07%,相比基于STOI的基线相对减少22.4%,显著提升了归一化互相关性能。
原文摘要 · Abstract (English)
We present LIWhiz, a non-intrusive lyric intelligibility prediction system submitted to the ICASSP 2026 Cadenza Challenge. LIWhiz leverages Whisper for robust feature extraction and a trainable back-end for score prediction. Tested on the Cadenza Lyric Intelligibility Prediction (CLIP) evaluation set, LIWhiz achieves a root mean square error (RMSE) of 27.07%, a 22.4% relative RMSE reduction over the STOI-based baseline, yielding a substantial improvement in normalized cross-correlation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。