arXiv:2608.07285eess.AScs.AI2026-08中稿 · International Soci…

首次将音乐中AI成分比例量化为0-1连续值,突破了传统二元判断局限。

How Much AI Is in This Track? Quantifying the Proportion of AI-Generated Stems in Hybrid Music Mixtures

  • 将AI音乐检测从二分类改为回归,直接预测混合音频中AI成分占比
  • 在真实混音场景下,模型对AI比例估计误差仅0.076(MAE)
  • 发现鼓和吉他更易被识别,而人声和贝斯更难检测,与频谱特征相关

AI生成音乐正越来越多地以音轨形式应用于制作流程,如合成鼓点、贝斯线或人声。然而现有检测系统仍为二元判断,将作品视为完全由AI或人类创作。本文将该问题重新定义为回归任务,基于连续的AI能量比α∈[0,1]进行建模。研究提出一种方法:利用多轨音乐数据集,构建包含已知比例的人类演奏与神经音频编解码器重建的AI音轨混合样本。实验表明,原本在纯数据上达到99%准确率的卷积神经网络,在处理混合音轨时输出随AI能量上升而变化,但存在噪声且校准不准。分析显示,检测敏感性依赖于乐器类型,鼓与吉他因编码伪影显著更易识别,而人声与贝斯则较难探测。在此基础上,训练一个用于α回归的同类CNN模型,在独立测试集上取得MAE=0.076,R²=0.85。结果表明,该回归框架是应对现实音乐生产中复杂混合场景的初步有效方案。

原文摘要 · Abstract (English)

AI-generated music is increasingly used at the stem level, with producers integrating synthetic drums, basslines, or vocals alongside human-performed instruments. However, current AI music detection systems are binary, treating tracks as either fully AI or fully human. In this paper, we reformulate AI music detection as a regression problem on a continuous AI energy ratio, alpha in [0, 1]. We propose a methodology that leverages a multi-track music dataset to assemble mixtures of human-performed and AI-reconstructed stems (obtained using a neural audio codec) with known proportions of each content type. Using this approach, we first show that a CNN-based model trained on fully AI-generated or human-performed tracks, which achieves >99% accuracy as a binary detector, when faced with mixed content, yields an output that rises with the AI stems' energy contribution, acting as a noisy and miscalibrated estimator. Our analysis of the influence of different stems shows that detection sensitivity depends on the instrument and reflects its frequency content: drums and guitar carry strong codec-artifact signatures, while vocals and bass are less detectable. Based on these insights, we train a similar CNN-based model for regression of alpha, achieving MAE = 0.076 and R^2 = 0.85 on held-out mixtures from the same pipeline. These results suggest that the regression formulation is an initial promising step towards AI-music detection in realistic music production workflows.

AI音乐音轨检测量化分析混合音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。