无需查询即可攻击私有脑电模型,利用公开编码器生成高迁移性对抗样本。
SW-ProxyCE: Zero-Query Adversarial Transfer from Public EEG Encoders to Private Downstream Models

- 基于小规模标注数据重构任务级决策几何,无需训练代理模型。
- 在三种脑电任务中成功攻击私有下游模型,迁移成功率超基线方法。
- 揭示脑电基础模型虽具强泛化能力但缺乏对抗鲁棒性,适合安全研究者关注。
脑电图(EEG)基础模型通过从大规模异构神经记录中学习可复用表征,成为脑电解码的新兴范式。然而,公开发布的EEG基础编码器虽推动下游发展,也带来新安全风险:公开表征可能使私有下游模型易受攻击。本文研究在公开编码器与私有下游模型场景下的对抗迁移攻击,攻击者拥有白盒访问权的发布编码器及少量任务匹配的标注参考集,但无法访问或查询目标模型的参数、输出或梯度。我们提出无查询的任务感知攻击框架SW-ProxyCE,通过收缩-去相关类原型从少量标注数据恢复任务级决策几何,实现无需额外训练的可迁移对抗样本生成。在三个EEG任务中,使用三种通用基础编码器和一个领域特定预训练编码器,覆盖跨被试与同被试场景下的线性探测与全微调下游模型进行评估。结果表明,仅凭公共编码器与有限标注参考集生成的对抗样本可有效迁移至不可访问的下游模型。SW-ProxyCE始终优于任务无关的表示偏移攻击,揭示了EEG基础模型虽具强迁移性,却不具备对抗鲁棒性。代码将开源于GitHub。
原文摘要 · Abstract (English)
Electroencephalography (EEG) foundation models have recently emerged as a promising paradigm for EEG decoding by learning reusable representations from large-scale heterogeneous neural recordings. However, the open release of EEG foundation encoders, while facilitating downstream developments, also introduces a previously unexplored security risk: publicly available representations may make private downstream models vulnerable. This paper investigates adversarial transfer attacks in EEG foundation model deployment in a public-encoder and private-downstream setting, where attackers have white-box access to a released encoder and a small task-matched labeled reference set, but no access or query to victim parameters, outputs, or gradients. We propose Shrinkage-Whitened Proxy Cross-Entropy (SW-ProxyCE), a query-free task-aware attack framework that recovers task-level decision geometry from a small labeled reference set through shrinkage-whitened class prototypes, enabling transferable adversarial generation without training an additional surrogate classifier. We evaluated SW-ProxyCE across three EEG tasks using three general-purpose foundation encoders and a paradigm-specific pre-trained encoder, covering both linear-probing and full-fine-tuning downstream models in cross-subject and within-subject scenarios. Results demonstrated that adversarial examples generated from the public encoder and limited labeled references can effectively transfer to inaccessible downstream models. SW-ProxyCE consistently outperformed task-agnostic representation-shift attacks, revealing that the strong transferability of EEG foundation models does not necessarily lead to adversarial robustness. Our code will be available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。