测试音频增强下假音乐检测模型的鲁棒性,发现轻微增强就大幅降低准确率。
Evaluating Fake Music Detection Performance Under Audio Augmentations
- 构建真实与合成音乐数据集,施加多种音频变换测试模型。
- 轻度音频增强使顶尖检测模型准确率显著下降。
- 提醒检测系统在真实场景中可能失效,适合安全与可信研究者参考。
随着生成音频模型的快速发展,区分真人创作与生成音乐变得愈发困难。为此,已有假音乐检测模型被提出。本文研究此类系统在音频增强下的鲁棒性。我们构建了一个包含真实与多系统生成音乐的数据集,并施加多种音频变换,分析其对分类准确率的影响。实验测试了一款近期最先进的音乐深度伪造检测模型在音频增强下的表现。结果表明,即使引入轻微增强,模型性能也显著下降。
原文摘要 · Abstract (English)
With the rapid advancement of generative audio models, distinguishing between human-composed and generated music is becoming increasingly challenging. As a response, models for detecting fake music have been proposed. In this work, we explore the robustness of such systems under audio augmentations. To evaluate model generalization, we constructed a dataset consisting of both real and synthetic music generated using several systems. We then apply a range of audio transformations and analyze how they affect classification accuracy. We test the performance of a recent state-of-the-art musical deepfake detection model in the presence of audio augmentations. The performance of the model decreases significantly even with the introduction of light augmentations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。