用深度学习检测AI生成的音乐,应对生成音乐带来的版权与伦理挑战。
Detecting Musical Deepfakes
- 基于梅尔频谱图训练卷积神经网络识别音乐深伪
- 在变速变调条件下仍保持较高检测准确率
- 为保护创作者和推动AI音乐良性发展提供技术支撑
Text-to-Music(TTM)平台的兴起使音乐创作更加普及,用户可轻松生成高质量作品。然而这也给音乐人和整个产业带来新挑战。本研究利用FakeMusicCaps数据集,通过分类音频是否为深度伪造来检测AI生成歌曲。为模拟真实对抗环境,对数据进行了节奏拉伸和音高偏移处理。从处理后的音频生成梅尔频谱图,并用于训练和评估卷积神经网络。除了技术成果,本文还探讨了TTM平台的伦理与社会影响,强调设计良好的检测系统对保护艺术家和释放生成式AI在音乐中的积极潜力至关重要。
原文摘要 · Abstract (English)
The proliferation of Text-to-Music (TTM) platforms has democratized music creation, enabling users to effortlessly generate high-quality compositions. However, this innovation also presents new challenges to musicians and the broader music industry. This study investigates the detection of AI-generated songs using the FakeMusicCaps dataset by classifying audio as either deepfake or human. To simulate real-world adversarial conditions, tempo stretching and pitch shifting were applied to the dataset. Mel spectrograms were generated from the modified audio, then used to train and evaluate a convolutional neural network. In addition to presenting technical results, this work explores the ethical and societal implications of TTM platforms, arguing that carefully designed detection systems are essential to both protecting artists and unlocking the positive potential of generative AI in music.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。