发现400M图片数据集中藏有数千张私密孕检超声图,存在身份泄露风险。
Privacy in Image Datasets: A Case Study on Pregnancy Ultrasounds
- 用CLIP相似度检索,系统筛查大规模图像数据集
- 检出数千条可识别身份的敏感信息,如姓名、地点
- 提醒数据集构建需加强隐私保护,尤其对医疗影像
生成模型兴起推动互联网大规模图像数据集使用,但常缺乏数据筛选。本文以孕检超声图为例,研究其在LAION-400M数据集中是否存在。通过CLIP嵌入相似性分析,我们成功检索到含孕检超声图像的样本,并检测出数千条高风险个人信息,包括姓名与地理位置。这些信息可能被用于重新识别或冒用身份。研究呼吁改进数据集构建流程,强化数据隐私保护与伦理规范。
原文摘要 · Abstract (English)
The rise of generative models has led to increased use of large-scale datasets collected from the internet, often with minimal or no data curation. This raises concerns about the inclusion of sensitive or private information. In this work, we explore the presence of pregnancy ultrasound images, which contain sensitive personal information and are often shared online. Through a systematic examination of LAION-400M dataset using CLIP embedding similarity, we retrieve images containing pregnancy ultrasound and detect thousands of entities of private information such as names and locations. Our findings reveal that multiple images have high-risk information that could enable re-identification or impersonation. We conclude with recommended practices for dataset curation, data privacy, and ethical use of public image datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。