arXiv:2508.08353cs.CYcs.AI2025-08被引 2

探讨合成数据是否属于欧盟隐私法中的个人数据

Processing of synthetic data in AI development for healthcare and the definition of personal data in EU law

  • 通过法律分析与实证研究,评估合成数据的匿名性
  • 发现合成数据在特定条件下可能不构成个人数据
  • 呼吁明确法规以推动医疗AI创新

人工智能有望变革医疗,但需访问健康数据。通过机器学习模型训练生成的合成数据可在保护隐私的同时实现数据共享。然而,通用数据保护条例(GDPR)在实践中的不确定性带来了行政负担,限制了合成数据的应用。本文通过系统分析相关法律文件及实证研究,探讨合成数据在GDPR下是否应被认定为个人数据。研究通过生成合成数据并模拟推断攻击,检验残留识别风险,挑战了对技术识别风险的普遍认知。结果表明,合成数据在特定条件下可能已匿名,但对‘合理可能风险’的界定仍存在模糊性。为促进创新,研究呼吁制定更清晰的法规,在隐私保护与医疗AI发展之间取得平衡。

原文摘要 · Abstract (English)

Artificial intelligence (AI) has the potential to transform healthcare, but it requires access to health data. Synthetic data that is generated through machine learning models trained on real data, offers a way to share data while preserving privacy. However, uncertainties in the practical application of the General Data Protection Regulation (GDPR) create an administrative burden, limiting the benefits of synthetic data. Through a systematic analysis of relevant legal sources and an empirical study, this article explores whether synthetic data should be classified as personal data under the GDPR. The study investigates the residual identification risk through generating synthetic data and simulating inference attacks, challenging common perceptions of technical identification risk. The findings suggest synthetic data is likely anonymous, depending on certain factors, but highlights uncertainties about what constitutes reasonably likely risk. To promote innovation, the study calls for clearer regulations to balance privacy protection with the advancement of AI in healthcare.

合成数据GDPR医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。