用多模态大模型评估歌声质感,摆脱参考曲与单一评分。
Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
- 构建四维专家标注数据集Sing-MD,涵盖气息、音色等维度。
- 提出VocalVerse模型,高效分析完整歌曲的长期声学特征。
- 设计人机协同排名基准,更贴近真实听感评价需求。
自动歌声评估对教育和娱乐至关重要。现有系统存在两大局限:依赖参考曲目,抑制创作表达;将复杂表演简化为仅基于音高和节奏的非诊断性分数。我们主张从判别式转向描述式评估,构建无参考、多维度的完整评估体系。首先,引入Sing-MD,一个由专家在四个维度(呼吸控制、音色质量、情感表达、发声技巧)上标注的大规模数据集。分析发现专家标注存在显著不一致性,挑战传统以准确率为指标的有效性。其次,针对多模态大语言模型(MLLM)处理长歌曲时的内存限制,提出VocalVerse——一种轻量级声学编码器与混合架构结合的高效方案,可建模全局性能特征与长期依赖关系。第三,为克服自动度量不足,建立H-TPR(人机协同分层感知排序)基准,评估模型生成符合听觉感知排序的能力,而非预测嘈杂的真实分数。
原文摘要 · Abstract (English)
Automated singing assessment is crucial for education and entertainment. However, existing systems face two fundamental limitations: reliance on reference tracks, which stifles creative expression, and the simplification of complex performances into non-diagnostic scores based solely on pitch and rhythm. We advocate for a shift from discriminative to descriptive evaluation, creating a complete ecosystem for reference-free, multi-dimensional assessment. First, we introduce Sing-MD, a large-scale dataset annotated by experts across four dimensions: breath control, timbre quality, emotional expression, and vocal technique. Our analysis reveals significant annotation inconsistencies among experts, challenging the validity of traditional accuracy-based metrics. Second, addressing the memory limitations of Multimodal Large Language Models (MLLMs) in analyzing full-length songs, we propose VocalVerse. This efficient hybrid architecture leverages a lightweight acoustic encoder to model global performance features and long-term dependencies. Third, to address automated metric shortcomings, we establish the H-TPR (Human-in-the-loop Tiered Perceptual Ranking) benchmark, which evaluates a model's ability to generate perceptually valid rankings rather than predicting noisy ground-truth scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。