攻击者可注入少量伪造歌词,让音乐生成系统偏离用户意图。
Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation

- 用双层策略在歌词中埋入低频声学特征,隐蔽操控生成方向。
- 仅用10个毒化歌词就使生成结果与攻击目标接近度提升47%。
- 适合关注创意AI安全性的研究人员和开发者参考。
检索增强型文本到音乐(TTM)系统通过从音乐歌词数据集检索内容来补全用户提示的不足,这使其对音乐知识库的完整性产生依赖。本文表明,攻击者可通过注入少量精心设计的音乐歌词,导致系统检索出恶意歌词,从而在不修改用户提示、检索器或生成器的前提下,使生成结果偏离用户原意。为此,我们提出一种双层歌词污染策略:在保持高层检索锚点不变的同时,嵌入低层声学描述符,以引导提示补全和下游音乐生成朝向攻击者指定的目标意图。在MusicCaps知识库、CLAP检索器和MusicGen生成流水线的实验中,中毒生成结果与攻击目标的相似度显著提升(平均+47%),同时仍与原始用户查询保持较高相关性。该研究揭示了检索增强型创意AI系统面临的真实完整性风险。演示链接:https://yizhu-wen.github.io/Mental-Damage/
原文摘要 · Abstract (English)
Retrieval-augmented text-to-music (TTM) systems augment underspecified user prompts using captions retrieved from a music caption dataset. This design introduces an integrity dependency on the music knowledge database. We show that an attacker can poison the database by injecting a small number of crafted music captions, causing the system to retrieve malicious captions that bias prompt augmentation and steer generation away from the user's intended function, without modifying the user prompt, retriever, or generator. To achieve the music caption poisoning attack, we propose a dual-layer caption poisoning strategy that preserves high-level retrieval anchors while injecting low-level acoustic descriptors to steer prompt augmentation and downstream music generation toward an attacker-chosen target intent. In a MusicCaps knowledge database, CLAP retriever, and MusicGen pipeline, poisoned generations move substantially closer to the attacker's target, while remaining comparably aligned with the original user query. These results expose a practical integrity risk for retrieval-augmented creative AI systems. Our demo can be found at: https://yizhu-wen.github.io/Mental-Damage/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。