arXiv:2602.02629cs.CRcs.AI2026-02被引 1

用区块链+数字身份保护医疗数据联邦学习,防伪身份攻击且性能不降。

Trustworthy Blockchain-based Federated Learning for Electronic Health Records: Securing Participant Identity with Decentralized Identifiers and Verifiable Credentials

  • 用去中心化标识和可验证凭证实现医疗实体身份认证
  • 100%抵御伪身份攻击,模型性能AUC达0.954、召回率0.890
  • 适合需要合规协作的医院与研究机构,成本仅约18美元/百轮

医疗数据数字化催生了海量电子健康记录(EHR),为训练AI模型提供了巨大机遇。但GDPR和HIPAA等严格隐私法规导致数据孤岛,阻碍集中式训练。联邦学习(FL)作为无需共享原始数据即可协作训练的方案受到关注。然而,现有方法仍易受投毒攻击和Sybil攻击影响,即恶意参与者通过伪造身份污染全局模型。尽管部分工作引入区块链以提升可审计性,但多依赖概率声誉系统而非强加密身份验证。本文提出可信区块链联邦学习(TBFL)框架,融合自我主权身份(SSI)标准,利用去中心化标识符(DIDs)和可验证凭证(VCs)确保仅经认证的医疗机构参与模型训练。基于MIMIC-IV数据集的全面评估表明,将信任锚定于密码学身份验证而非行为模式,显著降低安全风险并保持临床实用性。实验显示,该框架成功防御100%的Sybil攻击,预测性能优异(AUC=0.954,Recall=0.890),计算开销极小(<0.12%)。整体方案构建了一个安全、可扩展且经济可行的跨机构医疗数据协作生态,100轮训练总成本约为18美元。

原文摘要 · Abstract (English)

The digitization of healthcare has generated massive volumes of Electronic Health Records (EHRs), offering unprecedented opportunities for training Artificial Intelligence (AI) models. However, stringent privacy regulations such as GDPR and HIPAA have created data silos that prevent centralized training. Federated Learning (FL) has emerged as a promising solution that enables collaborative model training without sharing raw patient data. Despite its potential, FL remains vulnerable to poisoning and Sybil attacks, in which malicious participants corrupt the global model or infiltrate the network using fake identities. While recent approaches integrate Blockchain technology for auditability, they predominantly rely on probabilistic reputation systems rather than robust cryptographic identity verification. This paper proposes a Trustworthy Blockchain-based Federated Learning (TBFL) framework integrating Self-Sovereign Identity (SSI) standards. By leveraging Decentralized Identifiers (DIDs) and Verifiable Credentials (VCs), our architecture ensures only authenticated healthcare entities contribute to the global model. Through comprehensive evaluation using the MIMIC-IV dataset, we demonstrate that anchoring trust in cryptographic identity verification rather than behavioral patterns significantly mitigates security risks while maintaining clinical utility. Our results show the framework successfully neutralizes 100% of Sybil attacks, achieves robust predictive performance (AUC = 0.954, Recall = 0.890), and introduces negligible computational overhead (<0.12%). The approach provides a secure, scalable, and economically viable ecosystem for inter-institutional health data collaboration, with total operational costs of approximately $18 for 100 training rounds across multiple institutions.

联邦学习区块链医疗AI身份认证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。