arXiv:2506.15981cs.CLcs.AI2025-06ACL被引 8

融合音频与自动转录歌词,提升AI生成歌曲检测的鲁棒性。

Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion

  • 采用多视角融合架构,结合音频语音特征与自动生成的歌词。
  • 在多种音频扰动下仍保持高检测准确率,优于现有方法。
  • 适合版权方、音乐平台等实际场景使用,无需高质量歌词文本。

基于AI的音乐生成工具快速发展,正重塑音乐产业,但也给创作者、版权持有者和平台带来挑战。现有检测方法依赖音频或歌词,存在局限:音频基方法泛化能力差且易受噪声干扰;歌词基方法需清晰准确的歌词文本,实践中难以获得。为此,我们提出一种新颖的、贴近实际应用的方法:一个多模态、模块化晚融合管道,结合自动转录的演唱歌词与捕捉歌词相关信息的语音特征。通过直接从音频中提取歌词信息,该方法增强了鲁棒性,降低了对低层伪影的敏感度,并提升了实际可用性。实验表明,所提出的DE-detect方法在性能上超越现有歌词基检测器,同时对音频扰动更具鲁棒性。因此,它为真实场景中的AI生成音乐检测提供了有效且稳健的解决方案。代码已开源:https://github.com/deezer/robust-AI-lyrics-detection。

原文摘要 · Abstract (English)

The rapid advancement of AI-based music generation tools is revolutionizing the music industry but also posing challenges to artists, copyright holders, and providers alike. This necessitates reliable methods for detecting such AI-generated content. However, existing detectors, relying on either audio or lyrics, face key practical limitations: audio-based detectors fail to generalize to new or unseen generators and are vulnerable to audio perturbations; lyrics-based methods require cleanly formatted and accurate lyrics, unavailable in practice. To overcome these limitations, we propose a novel, practically grounded approach: a multimodal, modular late-fusion pipeline that combines automatically transcribed sung lyrics and speech features capturing lyrics-related information within the audio. By relying on lyrical aspects directly from audio, our method enhances robustness, mitigates susceptibility to low-level artifacts, and enables practical applicability. Experiments show that our method, DE-detect, outperforms existing lyrics-based detectors while also being more robust to audio perturbations. Thus, it offers an effective, robust solution for detecting AI-generated music in real-world scenarios. Our code is available at https://github.com/deezer/robust-AI-lyrics-detection.

AI音乐内容检测多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。