arXiv:2604.13610cs.CV2026-04

传统方法误把图像分辨率特征当语义差异,新方法用无监督聚类更真实评估数据集偏见。

What Are We Really Measuring? Rethinking Dataset Bias in Web-Scale Natural Image Collections via Unsupervised Semantic Clustering

论文配图:What Are We Really Measuring? Rethinking Dataset Bias in Web-Scale Natural Image Collections via Unsupervised Semantic Clustering
图 1 · 摘自论文原文
  • 用无监督聚类分析视觉特征,避开标签依赖的分类陷阱。
  • 在主流网络图像数据集中,分类准确率下降至接近随机水平。
  • 揭示旧方法因分辨率伪信号高估了语义偏见,适合关注数据公平性的研究者。

在计算机视觉中,衡量数据集偏见的常用方法是训练模型区分不同数据集,若分类准确率高,则认为存在显著语义差异。该方法假设标准图像增强能消除低层非语义线索,剩余性能反映真实语义分化。然而我们发现,在大规模自然图像集合中,这一假设存在根本缺陷:高分类准确率常由分辨率相关伪影驱动,这些伪影源于原始图像分辨率分布及重缩放时的插值效应,形成稳健的数据集特异性指纹,即使经过常规图像损坏仍存留。通过受控实验,我们证明模型在非语义的程序生成图像上也能实现强分类能力,说明其依赖的是表面线索。为解决此问题,我们重新审视数据集可分性概念,但采用无监督方法:直接基于基础视觉模型提取的语义丰富特征进行聚类,刻意绕开对数据集标签的监督分类。应用于主要网络规模数据集时,传统方法报告的高可分性几乎消失,聚类准确率降至接近随机水平。这表明,基于分类的评估方法系统性地夸大了语义偏见程度。

原文摘要 · Abstract (English)

In computer vision, a prevailing method for quantifying dataset bias is to train a model to distinguish between datasets. High classification accuracy is then interpreted as evidence of meaningful semantic differences. This approach assumes that standard image augmentations successfully suppress low-level, non-semantic cues, and that any remaining performance must therefore reflect true semantic divergence. We demonstrate that this fundamental assumption is flawed within the domain of large-scale natural image collections. High classification accuracy is often driven by resolution-based artifacts, which are structural fingerprints arising from native image resolution distributions and interpolation effects during resizing. These artifacts form robust, dataset-specific signatures that persist despite conventional image corruptions. Through controlled experiments, we show that models achieve strong dataset classification even on non-semantic, procedurally generated images, proving their reliance on superficial cues. To address this issue, we revisit this decades-old idea of dataset separability, but not with supervised classification. Instead, we introduce an unsupervised approach that measures true semantic separability. Our framework directly assesses semantic similarity by clustering semantically-rich features from foundational vision models, deliberately bypassing supervised classification on dataset labels. When applied to major web-scale datasets, the primary focus of this work, the high separability reported by supervised methods largely vanishes, with clustering accuracy dropping to near-chance levels. This reveals that conventional classification-based evaluation systematically overstates semantic bias by an overwhelming margin.

数据集偏见无监督学习视觉模型语义评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。