arXiv:2511.10913cs.SDcs.AI2025-11被引 1

研究大模型语音生成中内容滥用风险,提出有效绕过安全限制的方法。

Synthetic Voices, Real Threats: Evaluating Large Text-to-Speech Models in Generating Harmful Audio

  • 用语义混淆和音频通道注入法隐藏有害内容,突破安全拦截。
  • 在多个商业系统上将拒绝率降低,生成语音毒性显著提升。
  • 揭示平台检测与防护的漏洞,警示跨模态安全短板。

现代文本转语音(TTS)系统,尤其是基于大型音频语言模型(LALMs)的系统,能生成高保真语音,忠实还原输入文本并模仿指定说话人。以往滥用研究聚焦于说话人模仿,本文探索一种新的以内容为中心的威胁:利用TTS生成包含有害内容的语音。实现这一威胁面临两大挑战:(1) LALM的安全对齐机制常拒绝有害提示,但现有越狱攻击不适用于TTS,因这些系统设计为忠实朗读任何输入文本;(2) 实际部署管道常采用输入/输出过滤器,阻止有害文本和音频。我们提出HARMGEN,一套五种攻击方法,分为两类:第一类使用语义混淆技术(拼接、打乱)隐藏有害内容;第二类通过辅助音频通道(朗读、拼写、音素)注入有害内容,同时保持文本提示看似无害。在五个基于LALMs的商用TTS系统和三个涵盖两种语言的数据集上评估显示,我们的攻击显著降低拒绝率,提升生成语音的毒性。我们进一步评估了音频流平台的响应式对策和TTS厂商的主动防御。分析表明:深度伪造检测在高保真音频上表现不佳;响应式监管可被对抗扰动绕过;而主动防御可检测57%-93%的攻击。本工作揭示了TTS领域此前被忽视的内容滥用路径,强调需在训练与部署全链条建立鲁棒的跨模态防护。

原文摘要 · Abstract (English)

Modern text-to-speech (TTS) systems, particularly those built on Large Audio-Language Models (LALMs), generate high-fidelity speech that faithfully reproduces input text and mimics specified speaker identities. While prior misuse studies have focused on speaker impersonation, this work explores a distinct content-centric threat: exploiting TTS systems to produce speech containing harmful content. Realizing such threats poses two core challenges: (1) LALM safety alignment frequently rejects harmful prompts, yet existing jailbreak attacks are ill-suited for TTS because these systems are designed to faithfully vocalize any input text, and (2) real-world deployment pipelines often employ input/output filters that block harmful text and audio. We present HARMGEN, a suite of five attacks organized into two families that address these challenges. The first family employs semantic obfuscation techniques (Concat, Shuffle) that conceal harmful content within text. The second leverages audio-modality exploits (Read, Spell, Phoneme) that inject harmful content through auxiliary audio channels while maintaining benign textual prompts. Through evaluation across five commercial LALMs-based TTS systems and three datasets spanning two languages, we demonstrate that our attacks substantially reduce refusal rates and increase the toxicity of generated speech. We further assess both reactive countermeasures deployed by audio-streaming platforms and proactive defenses implemented by TTS providers. Our analysis reveals critical vulnerabilities: deepfake detectors underperform on high-fidelity audio; reactive moderation can be circumvented by adversarial perturbations; while proactive moderation detects 57-93% of attacks. Our work highlights a previously underexplored content-centric misuse vector for TTS and underscore the need for robust cross-modal safeguards throughout training and deployment.

语音生成安全风险对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。