分析语音模型中后门攻击如何在组件间传播及隐藏。
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
- 拆解语音模型各组件,定位后门传播路径。
- 后门能否留存取决于目标组件,部分组件可消除后门。
- 中毒样本与正常样本在共享嵌入中难以区分,挑战防御假设。
语音语言模型(SLMs)是多个独立组件协同工作的系统。尽管结构异构,现有研究多采用端到端方式,信息流动机制仍不清晰。本文聚焦后门攻击,首次证明后门可在整个语音模型中传播,导致所有任务均易受攻击。基于此,我们设计组件级分析方法,揭示各组件在后门学习中的作用。结果表明,后门的持续存在或消失高度依赖于目标组件。此外,我们研究了后门在共享多任务嵌入中的编码方式,发现中毒样本与正常样本在嵌入空间中无法直接分离,挑战了现有过滤类防御所依赖的可分性假设。研究强调,多模态流水线应被视为具有独特漏洞的复杂系统,而非单模态系统的简单延伸。
原文摘要 · Abstract (English)
Speech language models (SLMs) are systems of systems: independent components that unite to achieve a common goal. Despite their heterogeneous nature, SLMs are often studied end-to-end; how information flows through the pipeline remains obscure. We investigate this question through the lens of backdoor attacks. We first establish that backdoors can propagate through the SLM, leaving all tasks highly vulnerable. From this, we design a component analysis to discover the role each component takes in backdoor learning. We find that backdoor persistence or erasure is highly dependent on the targeted component. Beyond propagation, we examine how backdoors are encoded in shared multitask embeddings, showing that poisoned samples are not directly separable from benign ones, challenging a common separability assumption used in filtering defenses. Our findings emphasize the need to treat multimodal pipelines as intricate systems with unique vulnerabilities, not solely extensions of unimodal ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。