arXiv:2606.08663cs.SDeess.AS2026-06中稿 · ICML

研究AI音乐生成器迁移时的令牌空间差异,发现编码器类型影响检测效果。

Probing Token Spaces under Generator Shift in AI-Generated Music Detection

  • 用固定分类器比较不同音频令牌空间,控制下游模型一致。
  • Udio训练时X-Codec令牌最强,Suno-v3.5训练时MERT令牌更强。
  • 提示应将编码风格令牌空间作为检测迁移的核心考察维度。

AI生成音乐检测器在标准基准上表现稳健,但实际部署需面对训练中未见的生成器源。本文在 extsc{MoM-open} 上开展源受限评估,该数据集重构了 MoM-CLAM,以 FMA 和 MTG-Jamendo 替代不可分发的真实语料,同时保留伪造生成协议。为隔离表示层作用,引入 extsc{CoMoE}:一种紧凑固定分类器,可在保持下游架构与训练方式不变的前提下,比较异构音频令牌空间。实验表明,标准与真实源受限分割几乎饱和,而伪造源受限条件下暴露显著令牌空间差异:仅在 Udio 训练时,X-Codec 令牌表现最优;仅在 Suno-v3.5 训练时,MERT 派生令牌更优。结果表明,在生成器迁移场景下,应将编码器风格离散令牌空间作为首要实验轴线。代码与数据已公开于 https://github.com/MAAP-LAB/CoMoE。

原文摘要 · Abstract (English)

AI-generated music detectors can appear robust on standard benchmark splits, yet their deployments require transfer to generator sources absent during training. We study this problem with source-restricted evaluation on \textsc{MoM-open}, an open reconstruction of MoM-CLAM that replaces the non-redistributable real corpus with FMA and MTG-Jamendo while preserving the fake-generator protocol. To isolate the role of representation, we introduce \textsc{CoMoE}, a compact fixed classifier for comparing heterogeneous audio token spaces while keeping the downstream architecture and training recipe unchanged. Experiments show that standard and real-source-restricted splits are nearly saturated, whereas fake-source restriction exposes large differences between token spaces: X-Codec tokens are strongest when training on Udio alone, while MERT-derived tokens are stronger when training on Suno-v3.5 alone. These results suggest that codec-style discrete token spaces should be treated as a primary experimental axis under generator shift in AI-generated music detection. Our code and data are available at https://github.com/MAAP-LAB/CoMoE.

音乐生成检测令牌空间迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。