arXiv:2506.19108cs.SD2025-06中稿 · ISMIR 2025被引 19

揭示AI生成音乐的频谱伪影,提出高效可解释的检测方法

A Fourier Explanation of AI-music Artifacts

  • 通过傅里叶分析发现生成模型输出存在系统性频谱峰值
  • 在开源与商业AI音乐工具上验证,检测准确率超99%
  • 方法简单透明,适合监管与版权审核场景使用

生成式AI的快速发展重塑了音乐创作,但随之而来的版权争议、就业冲击与伦理问题引发广泛关注。尽管出现诸多AI内容检测服务,其机制仍高度不透明且由私企控制。本文研究合成内容的本质特征,分析生成模型中常见的反卷积模块,数学证明其输出必然产生系统性频率伪影——表现为微小但显著的谱峰。该现象源于特定模型架构,而非训练数据或权重。我们在多个开源模型及商业AI音乐生成器(如Suno、Udio)上验证理论,据此提出一种简单且可解释的检测准则。该方法虽结构简单,却在多个场景下达到与深度学习方法相当的检测精度,超过99%。

原文摘要 · Abstract (English)

The rapid rise of generative AI has transformed music creation, with millions of users engaging in AI-generated music. Despite its popularity, concerns regarding copyright infringement, job displacement, and ethical implications have led to growing scrutiny and legal challenges. In parallel, AI-detection services have emerged, yet these systems remain largely opaque and privately controlled, mirroring the very issues they aim to address. This paper explores the fundamental properties of synthetic content and how it can be detected. Specifically, we analyze deconvolution modules commonly used in generative models and mathematically prove that their outputs exhibit systematic frequency artifacts -- manifesting as small yet distinctive spectral peaks. This phenomenon, related to the well-known checkerboard artifact, is shown to be inherent to a chosen model architecture rather than a consequence of training data or model weights. We validate our theoretical findings through extensive experiments on open-source models, as well as commercial AI-music generators such as Suno and Udio. We use these insights to propose a simple and interpretable detection criterion for AI-generated music. Despite its simplicity, our method achieves detection accuracy on par with deep learning-based approaches, surpassing 99% accuracy on several scenarios.

AI音乐生成模型频谱分析可解释检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。