用超算放大数据不可学习性,提升医疗图像隐私保护
Scale-up Unlearnable Examples Learning with High-Performance Computing
- 在超算上分布式训练,大幅扩展样本不可学习的批量规模
- 不同数据集最优批大小各异,过大过小都影响保护效果
- 为医疗影像等敏感数据提供可定制的隐私防护方案
当前人工智能模型常保留用户交互数据,可能无意中包含敏感医疗信息。在放射科医生使用在线平台的AI诊断工具时,医学影像数据可能被未授权用于未来训练,引发隐私与知识产权风险。为此,新提出的不可学习样本(UE)方法旨在让数据难以被深度学习模型学习。其中,不可学习聚类(UC)方法在更大批量下表现更优,但受限于计算资源。本文利用Summit超算,通过分布式数据并行(DDP)训练,在Pets、MedMNist、Flowers、Flowers102等数据集上大规模扩展了UC学习。实验表明,批大小过大或过小均导致性能不稳定,影响准确性;但批大小与不可学习性的关系因数据集而异,需针对不同数据特征制定批大小策略。结果强调:选择合适批大小对防止学习、保障深度学习中的数据安全至关重要。
原文摘要 · Abstract (English)
Recent advancements in AI models are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources. To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE's unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。