arXiv:2603.23667cs.SDcs.AI2026-03被引 1

构建首个语义对齐的音乐深伪检测数据集,提升检测模型泛化能力。

Echoes: A semantically-aligned music deepfake detection dataset

  • 基于真实音频或歌曲描述生成伪造音乐,实现语义级对齐。
  • 在跨数据集测试中,现有模型在该数据集上性能下降超40%。
  • 适合关注音频安全、深度伪造检测的研究者使用。

我们提出Echoes,一个面向音乐深伪检测的新数据集,旨在训练和评估在现实且来源多样条件下的检测器。该数据集包含4,468首曲目(共131小时音频),覆盖流行、摇滚、电子等多种风格,并涵盖十种主流AI音乐生成系统产生的内容。为避免捷径学习并促进鲁棒泛化,数据集刻意设计为具有挑战性,确保伪造音频与真实参考音频在语义层面高度对齐,这一对齐通过直接以真实波形或歌曲描述作为生成条件实现。我们使用最先进的Wav2Vec2 XLS-R 2B表示,在跨数据集设置下评估Echoes,结果表明:(i) Echoes是当前最难的同类数据集;(ii) 在现有数据集上训练的检测器在Echoes上迁移效果差;(iii) 在Echoes上训练可获得最强泛化性能。这些发现表明,提供方多样性与语义对齐有助于学习更具迁移性的检测特征。

原文摘要 · Abstract (English)

We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provider-diverse conditions. Echoes comprises 4,468 tracks (131 hours of audio) spanning multiple genres (pop, rock, electronic), and includes content generated by ten popular AI music generation systems. To prevent shortcut learning and promote robust generalization, the dataset is deliberately constructed to be challenging, enforcing semantic-level alignment between spoofed audio and bona fide references. This alignment is achieved by conditioning generated audio samples directly on bona-fide waveforms or song descriptors. We evaluate Echoes in a cross-dataset setting against three existing AI-generated music datasets using state-of-the-art Wav2Vec2 XLS-R 2B representations. Results show that (i) Echoes is the hardest in-domain dataset; (ii) detectors trained on existing datasets transfer poorly to Echoes; (iii) training on Echoes yields the strongest generalization performance. These findings suggest that provider diversity and semantic alignment help learn more transferable detection cues.

音乐生成深伪检测数据集泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。