发现脑电基础模型即使冻结仍会泄露频谱特征,需联合审计才可靠。
Pretrained, Frozen, Still Leaking: Auditing Cross-Encoder Attribute Transfer in EEG Foundation Models
- 用跨编码器转移测试,验证冻结模型仍可泄露频谱属性。
- 六种方向中,跨主体解码置信区间下界均不低于0.081。
- 提出可部署的审计评分,比传统方法更敏感且不依赖模型头。
脑电基础模型发布常逐个评估单个端点:原始重建、成员推断、身份关联或下游头上的DP-SGD。我们对BIOT、LaBraM和EEGPT释放的嵌入表示,在四个端点上联合审计,发现单端点审计通过的模型仍存在频谱属性泄露。关键证据来自跨编码器转移审计:一个从单一冻结编码器学习的岭回归属性解码器,经线性桥接拟合后,能迁移至其他所有编码器的保留受试者测试集,所有六种方向中,受试者无关匹配对照组95%置信区间下界至少为0.081。我们证明充分条件:若两个编码器共享非平凡属性坐标投影重叠β,存在链式岭桥攻击者,中心增益下界为sqrt(β/(1+τ²)) - ε_br - ρ₀,反推得β ∈ [0.008, 0.198]。为将联合审计转化为可部署决策规则,引入审计端点分歧得分(AEDS),证明其正值的充分条件,并进行每单元自举校准;在八个匹配置信区间单元(BIOT/LaBraM/EEGPT在EEGMMI;LaBraM在Sleep-EDF、54通道LIMO、CHB-MIT儿科头皮EEG)中,AEDS均显著为正(p<0.001),而头部级Carlini LiRA成员推断审计仅达AUC 0.50–0.70。标准防御失效:维纳风格噪声感知自适应攻击、LiRA审计及所有保持效用的ε∈{4,8}下的DP-SGD,均未改变属性信道。贡献在于构建了将分散单端点防御整合为联合发布决策的审计框架,包含跨编码器桥定理与自适应攻击、LiRA及DP-SGD基线;该审计支持发布阻断,而非原始波形窃取或保留受试者身份恢复。
原文摘要 · Abstract (English)
EEG foundation-model releases are usually audited one endpoint at a time: raw-reconstruction, membership inference, identity linkage, or DP-SGD on the downstream head. We audit the same released embeddings under all four endpoints jointly, on BIOT, LaBraM, and EEGPT, and show that each single-endpoint audit clears releases that still leak spectral attributes. The decisive evidence is a cross-encoder transfer audit: a single ridge attribute decoder learned from one frozen encoder transfers, via a fitted linear bridge, to held-out-subject test splits of every other encoder, with subject-disjoint matched-control 95% CI lower bound at least 0.081 across all six BIOT/LaBraM/EEGPT directions. We prove a sufficient condition: two encoders sharing a nontrivial attribute-coordinate projector overlap beta admit a chained ridge bridge attacker with centered-gain lower bound sqrt(beta/(1+tau^2)) - eps_br - rho_0, and back-solve beta in [0.008, 0.198]. To turn the joint audit into a deployment-readable decision rule we introduce an audit-endpoint disagreement score (AEDS), prove sufficient conditions for its positivity, and bootstrap-calibrate it per cell; AEDS is positive in all eight matched-CI cells (BIOT/LaBraM/EEGPT on EEGMMI; LaBraM on Sleep-EDF, 54-channel LIMO, CHB-MIT pediatric scalp EEG) with p<0.001, while a head-level Carlini LiRA membership audit reaches AUC only 0.50-0.70. Standard defenses fail under audit: a Wiener-style noise-aware adaptive attacker, the LiRA audit, and DP-SGD at every utility-preserving epsilon in {4,8} leave the attribute channel essentially unchanged. The contribution is an audit framework that turns scattered single-endpoint defenses into a joint release decision, supported by a cross-encoder bridge theorem and adaptive-attacker, LiRA, and DP-SGD baselines; the audit licenses release-blocking, not raw-waveform exfiltration or held-out-subject identity recovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。