arXiv:2606.26451cs.SDcs.LG2026-06中稿 · Interspeech 2026

音乐感知框架让自动歌声评分更懂歌词与音准的结合

Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation

论文配图:Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation
图 1 · 摘自论文原文
  • 用多信号匹配识别歌词语义块,融合语义、词汇和发音对齐
  • 在多个数据集上与专家评分高度一致,表现稳定可靠
  • 适合需要精准评估演唱质量的音乐智能系统开发者

自动歌声质量评估(SQA)需同时判断歌词准确性与音乐契合度,但现有方法多仅依赖声学特征或歌词转写,难以全面评估。尤其在滑音、颤音和节奏弹性下,歌唱转写易出错,导致跨模态融合困难。为此,我们提出 MusicJudge,一种模态引导的自动化评估框架,通过分块对齐实现歌词与音高-节奏一致性的联合分析。该框架利用多信号匹配技术识别有意义的歌词片段,整合语义嵌入、词汇相似性和发音对齐。为提升歌唱音频转写效果,引入模态引导的 LoRA 进行语音识别模型微调。在多个数据集上的实验表明,MusicJudge 与人类专家评分具有强一致性,验证了其泛化能力。

原文摘要 · Abstract (English)

Automatic singing quality assessment (SQA) requires evaluating lyrical correctness and musical fidelity while handling expressive variations. However, existing systems largely rely on either acoustic cues or lyric transcriptions exclusively, limiting holistic performance evaluation. Furthermore, their integration is non-trivial due to challenges in robust singing transcription amid melisma, vibrato, and tempo elasticity. To this end, we propose MusicJudge, a modality-guided framework for automated SQA that performs block-aligned multimodal analysis by coupling lyric correctness with pitch-rhythm fidelity. It detects semantically meaningful lyric blocks using multi-signal matching that integrates semantic embeddings, lexical similarity, and phonetic alignment. To improve singing audio transcription, we introduce Modality-Guided LoRA for ASR fine-tuning. Experiments across datasets demonstrate strong agreement with human expert judgments and validate the generalizability of MusicJudge.

歌声评估多模态语音识别音乐智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。