arXiv:2501.10111cs.SDeess.AS2025-01中稿 · IEEE ICASSP 2025被引 18

首个可检测AI生成音乐的工具,准确率达99.8%。

AI-Generated Music Detection and its Challenges

  • 用真实与合成音频训练分类器,实现高精度识别。
  • 检测准确率高达99.8%,验证了可行性。
  • 揭示鲁棒性与泛化挑战,提醒谨慎应用。

生成模型新纪元下,检测人工生成内容变得至关重要。尤其在用户友好平台能在数秒内生成可信的分钟级合成音乐,对流媒体服务构成欺诈威胁,也对人类艺术家造成不公平竞争。本文展示通过真实音频与人工重构数据集训练分类器的可能性,并取得99.8%的惊人准确率。据我们所知,这是首个公开发布的AI音乐检测工具,将助力合成媒体监管。然而,基于其他领域伪造检测的长期研究,我们强调:高测试得分并非终点。本文揭示并讨论部署中可能存在的问题,如对音频篡改的鲁棒性、对未见生成模型的泛化能力。这部分为该领域未来研究指明方向,也为蓬勃发展的虚假内容检测市场敲响警钟。

原文摘要 · Abstract (English)

In the face of a new era of generative models, the detection of artificially generated content has become a matter of utmost importance. In particular, the ability to create credible minute-long synthetic music in a few seconds on user-friendly platforms poses a real threat of fraud on streaming services and unfair competition to human artists. This paper demonstrates the possibility (and surprising ease) of training classifiers on datasets comprising real audio and artificial reconstructions, achieving a convincing accuracy of 99.8%. To our knowledge, this marks the first publication of a AI-music detector, a tool that will help in the regulation of synthetic media. Nevertheless, informed by decades of literature on forgery detection in other fields, we stress that getting a good test score is not the end of the story. We expose and discuss several facets that could be problematic with such a deployed detector: robustness to audio manipulation, generalisation to unseen models. This second part acts as a position for future research steps in the field and a caveat to a flourishing market of artificial content checkers.

AI音乐内容检测生成模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。