arXiv:2602.06823cs.SDcs.AI2026-02中稿 · ICASSP 2026被引 4

首个面向广播场景的AI音乐检测数据集,解决短时弱音与语音掩蔽难题

AI-Generated Music Detection in Broadcast Monitoring

  • 构建3294段广播风格音频片段,模拟真实电视音轨时长与音量关系
  • 现有模型在背景音乐或短时片段下F1分数降至60%以下,性能显著下降
  • 为工业级广播监测提供新基准,推动抗干扰检测技术发展

AI音乐生成已达到难以与人类创作区分的水平。尽管已有检测方法,但多针对干净、完整的流媒体音乐设计,不适用于广播场景:广播中的音乐常以短片段出现,且被主导语音掩盖,现有检测器在此条件下失效。本文提出AI-OpenBMAT,首个专为广播场景设计的AI音乐检测数据集,包含3,294个一分钟音频片段(总计54.9小时),其时长分布与音量关系符合真实电视音频特征,融合真人制作的背景音乐与使用Suno v3.5生成的风格匹配续接。我们评估了CNN基线和SpectTTTra等先进模型在信噪比与持续时间鲁棒性上的表现,并在完整广播场景中进行测试。所有设置下,流媒体场景表现优异的模型在音乐处于背景或时长较短时性能大幅下降,F1分数低于60%。结果表明,语音掩蔽与短音乐片段是当前检测面临的关键挑战,同时确立了AI-OpenBMAT作为满足工业广播需求的检测器研发基准。

原文摘要 · Abstract (English)

AI music generators have advanced to the point where their outputs are often indistinguishable from human compositions. While detection methods have emerged, they are typically designed and validated in music streaming contexts with clean, full-length tracks. Broadcast audio, however, poses a different challenge: music appears as short excerpts, often masked by dominant speech, conditions under which existing detectors fail. In this work, we introduce AI-OpenBMAT, the first dataset tailored to broadcast-style AI-music detection. It contains 3,294 one-minute audio excerpts (54.9 hours) that follow the duration patterns and loudness relations of real television audio, combining human-made production music with stylistically matched continuations generated with Suno v3.5. We benchmark a CNN baseline and state-of-the-art SpectTTTra models to assess SNR and duration robustness, and evaluate on a full broadcast scenario. Across all settings, models that excel in streaming scenarios suffer substantial degradation, with F1-scores dropping below 60% when music is in the background or has a short duration. These results highlight speech masking and short music length as critical open challenges for AI music detection, and position AI-OpenBMAT as a benchmark for developing detectors capable of meeting industrial broadcast requirements.

AI音乐检测基准广播监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。