arXiv:2606.16612cs.SDcs.LG2026-06

用音乐内在特征检测AI生成歌曲,更准更鲁棒。

Beyond Artifacts: Towards Generalizable Synthetic Song Detection via Music-Intrinsic Features

论文配图:Beyond Artifacts: Towards Generalizable Synthetic Song Detection via Music-Intrinsic Features
图 1 · 摘自论文原文
  • 通过专家模型提取音乐内在特征,动态融合提升检测能力。
  • 在MUSIC8K数据集上相较最强基线F1提升18.5点。
  • 适合关注AI音乐检测、生成内容溯源的研究者。

AI音乐生成技术快速发展,亟需可靠的合成歌曲检测方法。现有方法多依赖低层伪影或固定特征假设,难以捕捉跨生成器的通用线索。为此,我们提出Sofia(基于音乐特征的合成歌曲检测框架),通过特征专用专家与自适应混合专家(MoE)模块,建模音乐内在属性。配置声乐、音频效果、全局结构特征及其组合,分析其独立与互补贡献。为全面评估,构建了MUSIC8K基准数据集,包含最新生成器及真实音频扰动。实验表明,Sofia从音乐内在特征中学习到生成器无关表征,在MUSIC8K-O上比最强基线F1提升18.5点,同时保持强鲁棒性。

原文摘要 · Abstract (English)

The rapid advancement of AI music generators highlights the urgent need for reliable Synthetic Song Detection (SSD). Existing SSD methods often rely on low-level artifacts or fixed feature assumptions, struggling to capture generator-agnostic cues. To address this, we propose Sofia (Synthetic-song detection framework via music features), a flexible framework that models music-intrinsic attributes via feature-specific experts and an adaptive Mixture-of-Experts (MoE) module. By configuring Sofia with representative Vocal, Audio-effect, Global structure features, and their combinations, we present their individual and complementary contributions. To comprehensively evaluate our framework, we further construct MUSIC8K, a challenging benchmark featuring lastest emerging generators and realistic audio perturbations. Experiments show that Sofia learns generator-agnostic representations from music-intrinsic features, improving the F1 score by 18.5 points over the strongest baseline on MUSIC8K-O while maintaining strong robustness.

AI音乐检测特征建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。