arXiv:2601.13647cs.SDcs.AI2026-01被引 1

提出融合双向注意力的模型,精准识别长时音乐中的AI生成内容。

Fusion Segment Transformer: Bi-Directional Attention Guided Fusion Network for AI-Generated Music Detection

  • 用多特征提取器分析短片段,再通过门控融合层整合内容与结构信息
  • 在SONICS和AIME数据集上达到当前最优性能,显著优于先前模型
  • 适合需要检测长音频中AI生成痕迹的研究者与版权保护应用

随着生成式AI技术的发展,任何人都能轻松创建并部署AI生成音乐,这加剧了版权与所有权问题的技术解决需求。现有研究主要聚焦于短音频检测,而对需建模长期结构与上下文的全音频检测仍缺乏充分探索。为此,我们提出改进版分段变换器——融合分段变换器(Fusion Segment Transformer)。延续前期工作,我们使用多种特征提取器从短音乐片段中提取内容嵌入;同时,通过引入门控融合层,有效整合内容与结构信息,以捕捉长时上下文。在SONICS与AIME数据集上的实验表明,该方法优于先前模型及近期基线,实现了AI生成音乐检测的最先进性能。

原文摘要 · Abstract (English)

With the rise of generative AI technology, anyone can now easily create and deploy AI-generated music, which has heightened the need for technical solutions to address copyright and ownership issues. While existing works mainly focused on short-audio, the challenge of full-audio detection, which requires modeling long-term structure and context, remains insufficiently explored. To address this, we propose an improved version of the Segment Transformer, termed the Fusion Segment Transformer. As in our previous work, we extract content embeddings from short music segments using diverse feature extractors. Furthermore, we enhance the architecture for full-audio AI-generated music detection by introducing a Gated Fusion Layer that effectively integrates content and structural information, enabling the capture of long-term context. Experiments on the SONICS and AIME datasets show that our approach outperforms the previous model and recent baselines, achieving state-of-the-art results in AI-generated music detection.

音乐生成AI检测长序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。