用Whisper嵌入实现可复现的歌词匹配,提升音乐检索可靠性
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
- 基于Whisper解码器嵌入构建可复现的歌词匹配流程
- 在标准数据集上性能媲美顶尖不可复现方法
- 支持多模态融合,适合音乐信息检索研究者
基于音频的歌词匹配可作为内容检索的有力替代方案,但现有方法常因缺乏可复现性与不一致的基线而受限。本文提出WEALY,一个完全可复现的流水线,利用Whisper解码器嵌入进行歌词匹配任务。WEALY建立了稳健透明的基线,并探索了融合文本与声学特征的多模态扩展。在标准数据集上的大量实验表明,WEALY性能可媲美缺乏可复现性的最先进方法。此外,我们还进行了消融研究,分析了语言鲁棒性、损失函数及嵌入策略的影响。本工作为未来研究提供了可靠基准,凸显了语音技术在音乐信息检索中的潜力。
原文摘要 · Abstract (English)
Audio-based lyrics matching can be an appealing alternative to other content-based retrieval approaches, but existing methods often suffer from limited reproducibility and inconsistent baselines. In this work, we introduce WEALY, a fully reproducible pipeline that leverages Whisper decoder embeddings for lyrics matching tasks. WEALY establishes robust and transparent baselines, while also exploring multimodal extensions that integrate textual and acoustic features. Through extensive experiments on standard datasets, we demonstrate that WEALY achieves a performance comparable to state-of-the-art methods that lack reproducibility. In addition, we provide ablation studies and analyses on language robustness, loss functions, and embedding strategies. This work contributes a reliable benchmark for future research, and underscores the potential of speech technologies for music information retrieval tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。