用语音合成伪造恶意音频,悄悄植入后门让情感识别系统出错
Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning

- 用文本转语音生成隐蔽声学触发器,可嵌入自然与合成语音中
- 仅需低比例污染数据,攻击成功率超90%,正常输入几乎不受影响
- 自监督模型易被攻陷,跨模型迁移性强,适合研究安全防御者
语音情感识别(SER)系统越来越多地依赖自监督声学表征,但其在训练阶段面临潜在攻击的威胁仍缺乏系统研究。本文首次系统性地探讨基于污染的后门攻击对SER的影响,重点聚焦由文本到语音(TTS)生成音频带来的威胁。我们提出一种隐蔽且低能耗的声学触发器,可无感知地嵌入自然与合成语音中,实现可扩展、一致性的污染。实验表明,即使在极低污染比例下,SER模型仍可被可靠攻陷,攻击成功率高,且对良性输入性能几乎无损。进一步发现,后门模式具有强跨模型迁移性,自监督表征尤其容易学习这些触发器。结果揭示,TTS技术极大降低了有效后门攻击的门槛,暴露出现代SER流水线中的关键漏洞,亟需针对性防御机制。
原文摘要 · Abstract (English)
Speech Emotion Recognition (SER) systems increasingly leverage self-supervised acoustic representations, yet their vulnerability to training-time attacks remains largely underexplored. This paper presents the first systematic study of poisoning-based backdoor attacks on SER, with a focus on threats enabled by text-to-speech (TTS) generated audio. We introduce a stealthy, low-energy acoustic trigger that can be embedded imperceptibly into both natural and synthetic speech, enabling scalable and consistent poisoning. Our experiments demonstrate that SER models can be reliably compromised with high attack success rates under low poisoning ratios, while maintaining near-clean performance on benign inputs. We further show that backdoor patterns exhibit strong cross-model transferability and that self-supervised representations are particularly susceptible to learning these triggers. These findings reveal that TTS technology dramatically lowers the barrier to effective backdoor attacks, exposing critical vulnerabilities in modern SER pipelines and motivating the urgent need for dedicated defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。