通过分析音乐片段结构,提升AI生成音乐的识别准确率。
Segment Transformer: AI-Generated Music Detection via Music Structural Analysis
- 用Transformer框架整合多种预训练模型提取音乐特征。
- 在短音频和完整音频上均达到高检测准确率。
- 适合音乐版权保护与AI内容溯源场景使用。
音频与音乐生成技术在音乐信息检索(MIR)领域取得了显著进展,但随之而来的是版权归属与作者身份不明的问题。当前难以明确区分一段音乐是人工智能生成还是人类创作。为解决这一挑战,本文通过分析音乐片段的结构模式,提升AI生成音乐(AIGM)检测的准确性。具体而言,为从短音频片段中提取音乐特征,我们集成多种预训练模型,包括自监督学习(SSL)模型或音频效果编码器,并将其纳入所提出的基于Transformer的框架。针对长音频,我们设计了分段Transformer,将音乐划分为多个片段,并学习片段间的相互关系。实验采用FakeMusicCaps和SONICS数据集,在短音频与全音频检测任务中均取得高准确率。结果表明,将分段级音乐特征融入长时序分析,可有效提升AIGM检测系统的性能与鲁棒性。
原文摘要 · Abstract (English)
Audio and music generation systems have been remarkably developed in the music information retrieval (MIR) research field. The advancement of these technologies raises copyright concerns, as ownership and authorship of AI-generated music (AIGM) remain unclear. Also, it can be difficult to determine whether a piece was generated by AI or composed by humans clearly. To address these challenges, we aim to improve the accuracy of AIGM detection by analyzing the structural patterns of music segments. Specifically, to extract musical features from short audio clips, we integrated various pre-trained models, including self-supervised learning (SSL) models or an audio effect encoder, each within our suggested transformer-based framework. Furthermore, for long audio, we developed a segment transformer that divides music into segments and learns inter-segment relationships. We used the FakeMusicCaps and SONICS datasets, achieving high accuracy in both the short-audio and full-audio detection experiments. These findings suggest that integrating segment-level musical features into long-range temporal analysis can effectively enhance both the performance and robustness of AIGM detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。