arXiv:2604.10424cs.LG2026-04被引 1

检测心电图基础模型训练数据参与隐私,发现即使不暴露原始数据也存在泄露风险。

Membership Inference Attacks Expose Participation Privacy in ECG Foundation Encoders

  • 针对自监督心电图编码器设计成员推断攻击,模拟三种真实攻击接口。
  • 小规模或机构特异性数据集的参与信息泄露严重,对比学习模型在嵌入空间中泄漏更明显。
  • 仅限制原始数据访问不足以保护隐私,需部署时开展针对性审计。

以自监督学习预训练的心电图基础编码器正被广泛跨任务、跨机构复用,常通过模型即服务接口暴露标量得分或潜在表示。尽管提升数据效率与泛化能力,却引发参与隐私问题:攻击者能否推断特定个体或群体是否参与过预训练?在连续健康场景中,训练参与本身可能暴露机构归属、研究注册或敏感健康背景。本文对现代自监督心电图基础编码器实施成员推断攻击审计,涵盖对比学习(SimCLR、TS2Vec)与掩码重建(基于CNN和Transformer的MAE)。评估三类现实攻击接口:(i)仅标量输出的黑盒访问;(ii)可聚合重复查询统计信息的自适应学习攻击者;(iii)可探测潜在表示几何结构的嵌入访问攻击者。采用以受试者为中心的协议,在跨数据集审计设置下固定假阳性率进行校准,发现参与泄露具有对象依赖性:小样本或机构特异性队列中泄露最显著,对比学习编码器在嵌入空间中甚至出现饱和泄露;而更大更多样本的数据集显著降低操作尾部风险。总体表明,仅限制原始信号或标签访问不足以保障参与隐私,亟需在连通健康系统中对可复用生物信号基础编码器进行部署感知的审计。

原文摘要 · Abstract (English)

Foundation-style ECG encoders pretrained with self-supervised learning are increasingly reused across tasks, institutions, and deployment contexts, often through model-as-a-service interfaces that expose scalar scores or latent representations. While such reuse improves data efficiency and generalization, it raises a participation privacy concern: can an adversary infer whether a specific individual or cohort contributed ECG data to pretraining, even when raw waveforms and diagnostic labels are never disclosed? In connected-health settings, training participation itself may reveal institutional affiliation, study enrollment, or sensitive health context. We present an implementation-grounded audit of membership inference attacks (MIAs) against modern self-supervised ECG foundation encoders, covering contrastive objectives (SimCLR, TS2Vec) and masked reconstruction objectives (CNN- and Transformer-based MAE). We evaluate three realistic attacker interfaces: (i) score-only black-box access to scalar outputs, (ii) adaptive learned attackers that aggregate subject-level statistics across repeated queries, and (iii) embedding-access attackers that probe latent representation geometry. Using a subject-centric protocol with window-to-subject aggregation and calibration at fixed false-positive rates under a cross-dataset auditing setting, we observe heterogeneous and objective-dependent participation leakage: leakage is most pronounced in small or institution-specific cohorts and, for contrastive encoders, can saturate in embedding space, while larger and more diverse datasets substantially attenuate operational tail risk. Overall, our results show that restricting access to raw signals or labels is insufficient to guarantee participation privacy, underscoring the need for deployment-aware auditing of reusable biosignal foundation encoders in connected-health systems.

隐私安全心电图成员推断基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。