构建160万图像的大规模语义分布外检测数据集,聚焦语义偏移挑战。
SOOD-ImageNet: a Large-Scale Dataset for Semantic Out-Of-Distribution Image Classification and Semantic Segmentation

- 利用视觉语言模型生成数据并人工校验,保证质量与规模
- 含56类共约160万张图像,覆盖图像分类与分割任务的分布外场景
- 专为研究语义漂移问题设计,适合评估真实世界中的模型泛化能力
计算机视觉中的分布外(OOD)检测是关键研究方向,相关基准对评估模型泛化能力及实际应用至关重要。然而现有文献中的OOD基准存在两大局限:(1) 常忽略语义偏移作为潜在挑战;(2) 规模远小于现代模型训练所用大规模数据集。为此,我们提出SOOD-ImageNet,一个包含约160万张图像、涵盖56个类别的新数据集,专为图像分类和语义分割等常见任务在分布外条件下的研究设计,特别关注语义偏移问题。通过结合现代视觉-语言模型的能力与精确的人工校验,我们确保了数据的可扩展性与高质量。通过对多种模型在SOOD-ImageNet上的广泛训练与评估,验证了其推动计算机视觉中OOD研究的巨大潜力。项目页面见:https://github.com/bach05/SOODImageNet.git。
原文摘要 · Abstract (English)
Out-of-Distribution (OOD) detection in computer vision is a crucial research area, with related benchmarks playing a vital role in assessing the generalizability of models and their applicability in real-world scenarios. However, existing OOD benchmarks in the literature suffer from two main limitations: (1) they often overlook semantic shift as a potential challenge, and (2) their scale is limited compared to the large datasets used to train modern models. To address these gaps, we introduce SOOD-ImageNet, a novel dataset comprising around 1.6M images across 56 classes, designed for common computer vision tasks such as image classification and semantic segmentation under OOD conditions, with a particular focus on the issue of semantic shift. We ensured the necessary scalability and quality by developing an innovative data engine that leverages the capabilities of modern vision-language models, complemented by accurate human checks. Through extensive training and evaluation of various models on SOOD-ImageNet, we showcase its potential to significantly advance OOD research in computer vision. The project page is available at https://github.com/bach05/SOODImageNet.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。