研究图像匿名化对图像检索的影响,发现原始数据训练的模型在匿名后表现更优。
Evaluating the Impact of Data Anonymization on Image Retrieval
- 用三类匿名方法、四档程度测试检索性能变化
- 匿名后原始数据训练模型仍保持最高检索相似度
- 为隐私合规的图像检索系统设计提供实证参考
随着《通用数据保护条例》等隐私法规的推行,视觉数据匿名化在机构中日益重要。然而,匿名化可能损害依赖视觉特征的计算机视觉系统性能,如基于内容的图像检索(CBIR)。尽管如此,匿名化对CBIR的影响尚未系统研究。本文基于正在由巴登-符腾堡州刑事警察局使用的DOKIQ项目(一种用于文档验证的人工智能系统),提出一个简单评估框架:匿名后的检索结果应尽可能接近原始数据的检索结果。我们使用两个公开数据集和内部DOKIQ数据集,系统评估了三种匿名化方法、四种匿名化程度及四种训练策略的影响,所有实验均基于最先进的DINOv2骨干网络。结果显示,经过原始数据训练的模型在匿名后仍能产生最相似的检索结果,表明存在明显的检索偏差。本研究为开发兼顾隐私与性能的CBIR系统提供了实用指导。
原文摘要 · Abstract (English)
With the growing importance of privacy regulations such as the General Data Protection Regulation, anonymizing visual data is becoming increasingly relevant across institutions. However, anonymization can negatively affect the performance of Computer Vision systems that rely on visual features, such as Content-Based Image Retrieval (CBIR). Despite this, the impact of anonymization on CBIR has not been systematically studied. This work addresses this gap, motivated by the DOKIQ project, an artificial intelligence-based system for document verification actively used by the State Criminal Police Office Baden-Württemberg. We propose a simple evaluation framework: retrieval results after anonymization should match those obtained before anonymization as closely as possible. To this end, we systematically assess the impact of anonymization using two public datasets and the internal DOKIQ dataset. Our experiments span three anonymization methods, four anonymization degrees, and four training strategies, all based on the state of the art backbone Self-Distillation with No Labels (DINO)v2. Our results reveal a pronounced retrieval bias in favor of models trained on original data, which produce the most similar retrievals after anonymization. The findings of this paper offer practical insights for developing privacy-compliant CBIR systems while preserving performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。