arXiv:2507.17937cs.SDcs.AI2025-07被引 1

用谐音替换歌词可绕过版权过滤,让生成音乐几乎一模一样。

Bob's Confetti: Phonetic Memorization Attacks in Music and Video Generation

  • 用同音异义词替换歌词,保留发音结构但避开关键词检测。
  • 生成音乐与原作相似度达91%,远超随机词(13.7%)和语义改写(42.2%)。
  • 该攻击可跨模态触发视频生成,适合关注内容安全与模型漏洞的研究者。

音乐与视频生成模型常依赖文本过滤防止复现受版权保护的内容。我们揭示了这一方法的重大漏洞,提出对抗性语音提示攻击(APT),利用模型对语音层面记忆的倾向——即模型将音素、押韵、重音、节奏等非词汇级声学模式与受版权保护内容绑定。APT 将标志性歌词替换为发音相似但语义无关的替代词(如“mom's spaghetti”变为“Bob's confetti”),在保持语音结构的同时规避词汇过滤。我们在领先歌词转歌曲模型(Suno、YuE)上评估 APT,涵盖英语与韩语的说唱、流行与K-pop歌曲。APT 在平均相似度上达到91%,远高于随机歌词(13.7%)和语义改写(42.2%)。嵌入分析证实:在 YuE 中,文本编码器对 APT 修改歌词与原文的余弦相似度为0.90,而 Sentence-BERT 语义相似度降至0.71,说明模型更依赖发音而非意义。该漏洞具有跨模态性——当仅以 APT 歌词提示 Veo 3 时,其能重建原视频场景,尽管提示中无视觉信息。此外,基于语音-语义特征的防御也失效,因 APT 提示的语义相似度高于正常改写。结果表明,亚词汇声学结构已成为跨模态检索密钥,使现有版权过滤机制系统性易受攻击。演示示例见 https://jrohsc.github.io/music_attack/。

原文摘要 · Abstract (English)

Generative AI systems for music and video commonly use text-based filters to prevent regurgitation of copyrighted material. We expose a significant vulnerability in this approach by introducing Adversarial PhoneTic Prompting (APT), a novel attack that bypasses these safeguards by exploiting phonetic memorization--the tendency of models to bind sub-lexical acoustic patterns (phonemes, rhyme, stress, cadence) to memorized copyrighted content. APT replaces iconic lyrics with homophonic but semantically unrelated alternatives (e.g., "mom's spaghetti" becomes "Bob's confetti"), preserving phonetic structure while evading lexical filters. We evaluate APT on leading lyrics-to-song models (Suno, YuE) across English and Korean songs spanning rap, pop, and K-pop. APT achieves 91% average similarity to copyrighted originals, versus 13.7% for random lyrics and 42.2% for semantic paraphrases. Embedding analysis confirms the mechanism: YuE's text encoder treats APT-modified lyrics as near-identical to originals (cosine similarity 0.90) while Sentence-BERT semantic similarity drops to 0.71, showing the model encodes phonetic structure over meaning. This vulnerability extends cross-modally--Veo 3 reconstructs visual scenes from original music videos when prompted with APT lyrics alone, despite no visual cues in the prompt. We further show that phonetic-semantic defense signatures fail, as APT prompts exhibit higher semantic similarity than benign paraphrases. Our findings reveal that sub-lexical acoustic structure acts as a cross-modal retrieval key, rendering current copyright filters systematically vulnerable. Demo examples are available at https://jrohsc.github.io/music_attack/.

内容安全语音攻击版权防护跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。