提出新方法评估完整歌曲的演唱质量,动态捕捉音质变化曲线。
Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves

- 分两阶段建模:先用教师模型生成伪标签预测片段质量,再聚合段落特征与评分
- 在多个数据集上提升13.95%的KTAU指标,优于现有最强基线
- 适合音乐AI、音频评估系统开发者,尤其关注长时演唱质量分析者
演唱质量评估(SQA)在多媒体应用和音乐AI中日益重要,但现有研究多聚焦短片段,难以适用于完整歌曲。与片段级评估不同,完整歌曲SQA需建模演唱质量在音频各段的演变规律及其对整体评价的影响。由于缺乏段级标注,直接为每个片段分配单一总分会忽略局部差异。为此,我们提出SongSQA框架:第一阶段使用预训练教师模型生成伪标签,训练段落评分预测器以实现无需人工标注的段级质量预测;第二阶段通过可学习的歌曲嵌入与自注意力机制,整合段落特征与预测评分,构建统一段落嵌入,动态聚合关键质量线索以生成整体质量预测,并输出随时间变化的段级质量曲线。实验表明,SongSQA在全歌SQA任务中表现优异,在所有数据集上持续提升各项指标,相较最强基线最高提升13.95%的KTAU值。
原文摘要 · Abstract (English)
Singing Quality Assessment (SQA) has become increasingly important for practical multimedia applications and Music AI systems, yet existing studies predominantly focus on short singing clips and remain insufficient for full-length songs. Unlike clip-level assessment, full-length song SQA requires modeling how singing quality varies across different audio segments and how these local variations influence the overall evaluation of vocal performance. Moreover, the scarcity of segment-level annotations makes effective supervision challenging, as directly assigning a single overall score label to every segment tends to treat different segment qualities as equivalent. To address these challenges, we propose SongSQA, a two-stage framework for full-length song SQA. In the first stage, a Segment Score Predictor is trained with pseudo labels generated by a pre-trained teacher model, enabling segment-level singing quality prediction without requiring manual segment annotations. In the second stage, a Song Quality Aggregator integrates segment features and predicted segment scores into unified segment embeddings, and employs a learnable song embedding together with self-attention to capture the connection between segment-level vocal performance and overall song quality. In this way, SongSQA dynamically aggregates critical quality cues across the song to produce a holistic quality prediction, while also generating a temporal segment-level quality curve. Experimental results demonstrate the effectiveness of SongSQA for full-length song SQA, achieving up to a 13.95% relative improvement in KTAU over the strongest baseline, while consistently improving other evaluation metrics across all datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。