arXiv:2607.19688cs.SD2026-07

提出五维诊断框架,精准定位AI翻唱中的音乐错误。

A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features

论文配图:A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features
图 1 · 摘自论文原文
  • 构建旋律、和声、调性等五个维度的诊断体系
  • 和声与编排错误率高达53%和47%,调性相对稳定
  • 适合音乐生成质量评估与缺陷分析的研究者

AI生成的翻唱常因局部音乐错误而失败,全局质量评分难以定位问题:例如人声轮廓仍可辨识,但伴奏使用了错误和声功能,或整体保持在调内但编排不完整。本文提出一个包含旋律音高、和声进行、调性一致性、风格一致性和编排/制作质量的五维诊断框架。基准数据集包含由6个系统从5首原曲生成的30首翻唱作品,配有专家严重程度评分及9个符号化或声学特征。结果显示,和声进行与编排错误率最高(分别为53%和47%),而调性一致性较好;有6首作品虽调性合理,却存在严重和声错误。大跳比例与旋律评分呈弱相关(Spearman rho = -0.429,未校正p = 0.018),但无特征相关性在多重检验下显著。可解释的百分位规则初步测试未能在16项维度比较中稳定超越固定多数基线。研究揭示:低层级符号化总结可暴露特定问题,但无法替代需要上下文感知的音乐判断。

原文摘要 · Abstract (English)

AI-generated covers often fail through local musical errors that a global quality score cannot locate: the vocal contour may remain recognizable while the accompaniment uses the wrong harmonic function, or the output may stay in key while the arrangement remains incomplete. We present a five-dimensional diagnostic framework covering melodic pitch, harmonic progression, key consistency, style consistency, and arrangement/production quality. The benchmark contains 30 covers generated from 5 source songs by 6 systems, with expert severity ratings and 9 symbolic or acoustic features. Harmonic progression and arrangement had the highest severe-error rates (53% and 47%), whereas key consistency was better preserved. Six covers combined acceptable key consistency with severe harmonic errors. Large-leap ratio had a nominal association with melodic ratings (Spearman rho = -0.429, uncorrected p = 0.018), but no feature correlation survived the nine-test multiplicity reference. An interpretable percentile-rule pilot likewise failed to outperform a fixed majority baseline reliably across 16 dimension-level comparisons. The results separate useful diagnostic evidence from dependable automatic scoring: low-level and symbolic summaries can expose particular symptoms, but they do not replace context-aware musical judgment.

音乐生成诊断框架和声分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。