arXiv:2512.00621cs.SDcs.AI2025-12中稿 · Transactions on Ma…被引 5

用双流对比学习检测合成音乐,准确率达92.5%

Melody or Machine: Detecting Synthetic Music with Dual-Stream Contrastive Learning

  • 双流架构分别处理人声与乐器,捕捉合成痕迹
  • 在13万首歌曲上实现0.925的F1分数
  • 适合需要高鲁棒性检测器的研究者

端到端AI音乐生成的快速发展对艺术真实性与版权构成日益严峻的威胁,亟需能跟上步伐的检测方法。现有模型如SpecTTTra在面对多样且快速演进的新生成器时表现不佳,尤其在分布外(OOD)内容上性能显著下降,暴露出泛化能力不足的关键缺口:亟需更具挑战性的基准和更鲁棒的检测架构。为此,我们首先提出Melody or Machine(MoM),一个包含超过13万首歌曲(6,665小时)的大规模新基准,是迄今最丰富的数据集,融合开源与闭源模型,并设计了专门用于促进真正泛化检测器发展的OOD测试集。同时引入CLAM,一种新颖的双流检测架构。我们假设:人声与乐器间细微的、机器导致的不一致,虽在混合信号中难以察觉,却是合成的有力线索。CLAM通过两个预训练音频编码器(MERT和Wave2Vec2)生成并行表示,再由可学习的交叉聚合模块建模其相互依赖关系。模型采用双损失目标训练:标准二元交叉熵损失用于分类,辅以对比三元损失,使模型学会区分一致与人为错配的流对组合,从而增强对合成伪影的敏感性,而不依赖于简单的特征对齐。CLAM在具有挑战性的MoM基准上达到新的最先进水平,F1得分为0.925。

原文摘要 · Abstract (English)

The rapid evolution of end-to-end AI music generation poses an escalating threat to artistic authenticity and copyright, demanding detection methods that can keep pace. While foundational, existing models like SpecTTTra falter when faced with the diverse and rapidly advancing ecosystem of new generators, exhibiting significant performance drops on out-of-distribution (OOD) content. This generalization failure highlights a critical gap: the need for more challenging benchmarks and more robust detection architectures. To address this, we first introduce Melody or Machine (MoM), a new large-scale benchmark of over 130,000 songs (6,665 hours). MoM is the most diverse dataset to date, built with a mix of open and closed-source models and a curated OOD test set designed specifically to foster the development of truly generalizable detectors. Alongside this benchmark, we introduce CLAM, a novel dual-stream detection architecture. We hypothesize that subtle, machine-induced inconsistencies between vocal and instrumental elements, often imperceptible in a mixed signal, offer a powerful tell-tale sign of synthesis. CLAM is designed to test this hypothesis by employing two distinct pre-trained audio encoders (MERT and Wave2Vec2) to create parallel representations of the audio. These representations are fused by a learnable cross-aggregation module that models their inter-dependencies. The model is trained with a dual-loss objective: a standard binary cross-entropy loss for classification, complemented by a contrastive triplet loss which trains the model to distinguish between coherent and artificially mismatched stream pairings, enhancing its sensitivity to synthetic artifacts without presuming a simple feature alignment. CLAM establishes a new state-of-the-art in synthetic music forensics. It achieves an F1 score of 0.925 on our challenging MoM benchmark.

音乐生成伪造检测双流网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。