arXiv:2508.17416cs.CV2025-08ICCV被引 8

发现视觉数据集普遍存在图像泄露,影响模型评估公正性

Data Leakage in Visual Datasets

  • 按模态、覆盖范围和程度分类分析图像泄露现象
  • 所有测试集均存在泄露,严重程度从明显到隐蔽不等
  • 适合关注模型评估可靠性与数据安全的研究者

我们分析了视觉数据集中存在的数据泄露问题。数据泄露指评估基准中的图像在训练阶段已被见过,从而损害模型评估的公平性。由于大规模数据集常来源于互联网,而许多计算机视觉基准公开可得,我们的研究聚焦于识别并分析这一现象。我们根据模态、覆盖范围和程度将视觉泄露分为不同类别。通过图像检索技术,我们明确证实所有被分析的数据集都存在某种形式的泄露,且各类泄露——从严重到细微——都会影响下游任务中模型评估的可靠性。

原文摘要 · Abstract (English)

We analyze data leakage in visual datasets. Data leakage refers to images in evaluation benchmarks that have been seen during training, compromising fair model evaluation. Given that large-scale datasets are often sourced from the internet, where many computer vision benchmarks are publicly available, our efforts are focused into identifying and studying this phenomenon. We characterize visual leakage into different types according to its modality, coverage, and degree. By applying image retrieval techniques, we unequivocally show that all the analyzed datasets present some form of leakage, and that all types of leakage, from severe instances to more subtle cases, compromise the reliability of model evaluation in downstream tasks.

数据泄露视觉模型评估可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。