构建首个针对艺术作品中嗅觉参考物的细粒度检测数据集
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
- 构建包含139类细粒度物体的4712张图像数据集
- 标注共38,116个物体,存在密集重叠与全图分布特征
- 面向艺术视觉遗产研究,挑战嗅觉感知与物体识别交叉任务
计算机视觉在人文学科中的实际应用需要算法具备对抗艺术抽象、边缘物体及细微类别差异的能力。现有数据集虽提供艺术作品的实例级标注,但普遍偏向图像中心,且类别细节不足。本文提出的ODOR数据集填补了这一空白,涵盖4712张图像中的38,116个物体级标注,覆盖139种细粒度类别。统计分析揭示其挑战性特征:类别细致、物体密集重叠、空间分布遍及整个画布。同时提供对象检测模型的基准分析,并通过多项次级研究凸显数据集难点。该数据集旨在推动艺术作品物体检测及更广泛的视觉文化遗产研究,激励探索物体识别与嗅觉感知的交叉领域。
原文摘要 · Abstract (English)
Real-world applications of computer vision in the humanities require algorithms to be robust against artistic abstraction, peripheral objects, and subtle differences between fine-grained target classes. Existing datasets provide instance-level annotations on artworks but are generally biased towards the image centre and limited with regard to detailed object classes. The proposed ODOR dataset fills this gap, offering 38,116 object-level annotations across 4712 images, spanning an extensive set of 139 fine-grained categories. Conducting a statistical analysis, we showcase challenging dataset properties, such as a detailed set of categories, dense and overlapping objects, and spatial distribution over the whole image canvas. Furthermore, we provide an extensive baseline analysis for object detection models and highlight the challenging properties of the dataset through a set of secondary studies. Inspiring further research on artwork object detection and broader visual cultural heritage studies, the dataset challenges researchers to explore the intersection of object recognition and smell perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。