系统梳理AI生成音乐检测方法,揭示其与音频深度伪造检测的迁移关系。
From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview
- 构建四层检测分类体系,按信号、特征、水印和语义一致性分层组织方法。
- 通过分层可迁移性分析,指出哪些检测组件可复用于AI音乐检测及其适用条件。
- 为研究者提供多维度分类框架,适合关注AI内容安全与检测技术的开发者。
随着人工智能技术不断发展,其在生成真实且符合上下文的内容方面已扩展至多个领域。音乐作为根植于人类文化的艺术形式和娱乐媒介,正越来越多地被AI参与创作。然而,缺乏监管的AI音乐生成(AIGM)工具使用可能对音乐产业、版权及艺术完整性带来负面影响,凸显出有效AIGM检测的重要性。本文系统综述现有AIGM检测方法,提出一个四层检测分类体系:信号级、特征级、水印、语义一致性,依据所利用的痕迹类型进行组织。基于更为成熟的音频深度伪造检测领域,进一步开展分层可迁移性分析,探讨哪些组件可或不可迁移至AIGM检测,并明确其适用条件。此外,通过多维分类,将代表性方法按输入模态、检测粒度、特征类型、模型类型、检测目标、鲁棒性设置和可解释性进行组织。最后讨论该领域的意义,并提出未来研究方向以应对持续存在的挑战。
原文摘要 · Abstract (English)
As Artificial Intelligence (AI) technologies continue to evolve, their use in generating realistic, contextually appropriate content has expanded into various domains. Music, an art form and medium for entertainment deeply rooted in human culture, is seeing an increased involvement of AI into its production. However, the unregulated use of AI music generation (AIGM) tools raises concerns about potential negative impacts on the music industry, copyright, and artistic integrity, underscoring the importance of effective AIGM detection. This paper provides a systematic overview of existing AIGM detection methods. We first establish a four-level detection taxonomy: signal-level, feature-level, watermark, and semantic consistency, organising methods according to the type of trace they exploit. Drawing on the more mature field of audio deepfake detection, we then present a stratified transferability analysis that examines which components may or may not transfer to AIGM detection, and under what conditions. A multi-dimensional classification further organises representative methods along input modality, detection granularity, feature type, model type, detection target, robustness setting, and interpretability. We conclude by discussing implications and proposing directions for future research to address ongoing challenges in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。