arXiv:2509.14149cs.CV2025-09中稿 · BMVC 2025

用分层抽象图像研究视觉表征能力,发现抽象程度影响模型性能。

An Exploratory Study on Abstract Images and Visual Representations Learned from Them

  • 构建分层抽象图像数据集HAID,从真实图像生成多级抽象图
  • 抽象图像上模型在分类、分割、检测任务中性能显著下降
  • 揭示抽象程度与语义信息保留的关系,适合视觉认知研究者

想象一个仅由基本几何形状构成的世界,你还能识别熟悉物体吗?最近研究表明,由基本形状构成的抽象图像确实能向深度学习模型传递视觉语义信息。然而,此类图像学到的表征性能仍不及传统栅格图像。本文探究这一性能差距的原因,并分析不同抽象层级下高层语义内容的保留程度。为此,我们提出分层抽象图像数据集(HAID),该数据集通过多级抽象从正常栅格图像生成抽象图像。我们在HAID上对常规视觉系统进行训练与评估,涵盖分类、分割和目标检测等任务,全面比较了栅格化图像与抽象图像表征的差异。同时讨论抽象图像是否可作为有效视觉语义表达形式,助力视觉任务。

原文摘要 · Abstract (English)

Imagine living in a world composed solely of primitive shapes, could you still recognise familiar objects? Recent studies have shown that abstract images-constructed by primitive shapes-can indeed convey visual semantic information to deep learning models. However, representations obtained from such images often fall short compared to those derived from traditional raster images. In this paper, we study the reasons behind this performance gap and investigate how much high-level semantic content can be captured at different abstraction levels. To this end, we introduce the Hierarchical Abstraction Image Dataset (HAID), a novel data collection that comprises abstract images generated from normal raster images at multiple levels of abstraction. We then train and evaluate conventional vision systems on HAID across various tasks including classification, segmentation, and object detection, providing a comprehensive study between rasterised and abstract image representations. We also discuss if the abstract image can be considered as a potentially effective format for conveying visual semantic information and contributing to vision tasks.

图像抽象视觉表征数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。