剖析大规模视觉数据集的偏见来源,揭示其内在特征差异。
Understanding Bias in Large-Scale Visual Datasets
- 通过多种图像变换提取语义、结构、颜色等信息,定位偏见根源。
- 发现不同数据集在边界、颜色、频率等维度存在显著差异。
- 结合自然语言生成描述,帮助研究者构建更均衡的数据集。
近期研究指出,大规模视觉数据集存在显著偏见,可被现代神经网络轻易分类。然而这些数据集的具体偏见形式尚不明确。本文提出一个框架,通过施加多种图像变换,提取并分析数据集中的语义、结构、边界、颜色和频域信息,评估各类信息对偏见的贡献度。进一步采用对象级分析分解语义偏见,并利用自然语言方法生成每类数据集的详细、开放式描述。本研究旨在帮助研究人员理解现有预训练数据集中的偏见,为未来构建更多样、更具代表性的数据集提供依据。项目页面与代码已公开于 http://boyazeng.github.io/understand_bias。
原文摘要 · Abstract (English)
A recent study has shown that large-scale visual datasets are very biased: they can be easily classified by modern neural networks. However, the concrete forms of bias among these datasets remain unclear. In this study, we propose a framework to identify the unique visual attributes distinguishing these datasets. Our approach applies various transformations to extract semantic, structural, boundary, color, and frequency information from datasets, and assess how much each type of information reflects their bias. We further decompose their semantic bias with object-level analysis, and leverage natural language methods to generate detailed, open-ended descriptions of each dataset's characteristics. Our work aims to help researchers understand the bias in existing large-scale pre-training datasets, and build more diverse and representative ones in the future. Our project page and code are available at http://boyazeng.github.io/understand_bias .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。