无需预先知道生成模型,就能识别AI音乐真伪与来源。
Finding the noise: Zero-shot AI Music Detection

- 用音频特征提取+非负矩阵分解,从音乐中捕捉生成痕迹。
- 零样本下对多种AI音乐生成器实现高精度分类与聚类。
- 适合快速应对新上线AI音乐服务的检测需求。
我们提出一种新方法,用于在检测器未知生成模型的情况下识别AI生成音乐(如新发布的Suno、Udio等服务)。由于2023年以来用户友好的AI音乐生成服务迅速增多且持续更新,亟需一种无需监督的检测方式以适应动态变化。本文研究两个任务:一是判别真实与合成音乐,采用单类分类思路,基于基准真实音乐识别异常;二是零样本多类识别,将真实与多种AI音乐混合数据进行无监督聚类,目标是形成纯净、一致的类别。我们结合先前提出的伪影提取方法,再应用非负矩阵分解及简单分类/聚类算法,在两项任务上均取得优异表现,证明该方法可有效监控大规模音乐库中来自多种新型生成模型的合成内容。
原文摘要 · Abstract (English)
We present a novel method for AI-generated music detection in scenarios where the models that generated the input samples are unknown to the detector (e.g., from a newly released service). Since 2023, there has been a multiplication of user-friendly AI-music generation services (e.g., Suno, Udio), along with regular updates and new features. There is thus a need to address synthetic content detection in an unsupervised way to adapt to this rapidly changing context. This angle has not been much studied in music yet. We propose to study two tasks. First, discriminating between real and synthetic music. This may be approached in a one-class manner, namely, using some baseline real music and trying to determine what falls outside. Second, zero-shot multi-class identification, which is more similar to an unsupervised clustering task on a mix of real and various AI-music generations, where the goal is to create coherent, high-purity clusters. We propose a combination of a previously proposed artifact-extraction method, on top of which we apply non-negative matrix factorization and simple classification and clustering methods. We achieve excellent performance on both tasks, showing that the proposed methods may be used to monitor large-scale catalogs that may receive AI-generated samples from various newly released generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。