高密度街景数据中,群体身份可被轻易推断,隐私风险远超人脸模糊。
Privacy of Groups in Dense Street Imagery
- 通过渗透测试发现,模糊处理后的人群仍可被识别群体身份。
- 在2500万张纽约街景图像中,成功推断出敏感群体归属。
- 提出群体隐私保护框架,适合数据研究人员参考。
空间和时间上密集的街景影像(DSI)数据集规模迅速增长,2024年仅企业就拥有约3万亿张公共街道图像。随着Lyft、Waymo等公司用DSI训练自动驾驶算法和分析事故,学术界也利用其开展城市分析研究。尽管DSI提供方通过模糊人脸和车牌来保护个人隐私,但这些措施无法应对更广泛的隐私威胁。本文发现,数据密度提升与人工智能进步使得从看似匿名的数据中推断敏感群体成员身份成为可能。我们通过渗透测试,在25,232,608张纽约市行车记录仪图像中验证了该风险。构建了DSI中可识别群体的分类体系,并基于情境完整性理论分析其隐私影响。最后提出面向研究者的可操作建议。
原文摘要 · Abstract (English)
Spatially and temporally dense street imagery (DSI) datasets have grown unbounded. In 2024, individual companies possessed around 3 trillion unique images of public streets. DSI data streams are only set to grow as companies like Lyft and Waymo use DSI to train autonomous vehicle algorithms and analyze collisions. Academic researchers leverage DSI to explore novel approaches to urban analysis. Despite good-faith efforts by DSI providers to protect individual privacy through blurring faces and license plates, these measures fail to address broader privacy concerns. In this work, we find that increased data density and advancements in artificial intelligence enable harmful group membership inferences from supposedly anonymized data. We perform a penetration test to demonstrate how easily sensitive group affiliations can be inferred from obfuscated pedestrians in 25,232,608 dashcam images taken in New York City. We develop a typology of identifiable groups within DSI and analyze privacy implications through the lens of contextual integrity. Finally, we discuss actionable recommendations for researchers working with data from DSI providers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。