arXiv:2601.12222cs.SDcs.MM2026-01中稿 · the 27th Internati…被引 1

用新模型评估音乐美感,更贴近人类感知。

Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling

  • 通过多轨注意力融合捕捉音乐复杂特征
  • 分层次建模评分概率,输出区间内回归结果
  • 适合评估生成音乐的审美质量,尤其擅长细节判断

音乐生成人工智能正快速扩充音乐内容,亟需自动化歌曲美学评估。然而现有研究多聚焦语音、音频或演唱质量,对歌曲美学关注不足。传统方法直接预测精确的平均意见分(MOS),难以捕捉人类对歌曲美感感知的细微差别。本文提出一种面向歌曲的美学评估框架,包含两个新模块:1)多轨注意力融合(MSAF)在混音-人声与混音-伴奏对之间建立双向交叉注意力,融合信息以捕捉复杂音乐特征;2)分层粒度感知区间聚合(HiGIA)学习多粒度评分概率分布,将其聚合为评分区间,并在区间内进行回归以得到最终得分。我们在两个全长歌曲数据集上进行了评估:SongEval数据集(AI生成)和一个内部美学数据集(人工创作),并与两种先进模型对比。结果表明,所提方法在多维度歌曲美学评估中表现更优。推理代码与模型检查点已公开于 https://github.com/yisan33/song-aesthetics-evaluation。

原文摘要 · Abstract (English)

Music generative artificial intelligence (AI) is rapidly expanding music content, necessitating automated song aesthetics evaluation. However, existing studies largely focus on speech, audio or singing quality, leaving song aesthetics underexplored. Moreover, conventional approaches often predict a precise Mean Opinion Score (MOS) value directly, which struggles to capture the nuances of human perception in song aesthetics evaluation. This paper proposes a song-oriented aesthetics evaluation framework, featuring two novel modules: 1) Multi-Stem Attention Fusion (MSAF) builds bidirectional cross-attention between mixture-vocal and mixture-accompaniment pairs, fusing them to capture complex musical features; 2) Hierarchical Granularity-Aware Interval Aggregation (HiGIA) learns multi-granularity score probability distributions, aggregates them into a score interval, and applies a regression within the interval to produce the final score. We evaluated on two datasets of full-length songs: SongEval dataset (AI-generated) and an internal aesthetics dataset (human-created), and compared with two state-of-the-art (SOTA) models. Results show that the proposed method achieves stronger performance for multi-dimensional song aesthetics evaluation. The inference code and checkpoint are publicly available at https://github.com/yisan33/song-aesthetics-evaluation.

音乐生成美学评估多模态评分建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。