arXiv:2603.28363cs.CV2026-03

用常识问答评估草图如何以少胜多,精准捕捉核心元素

SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering

  • 基于常识知识提取每类物体的关键视觉元素
  • 通过视觉问答模型量化草图对关键元素的保留程度
  • 首个带元素级标注的草图数据集,适合评估理解能力

草图是通过简化却有目的的笔触传达核心概念的视觉抽象形式,省略冗余细节。尽管表现力强,但量化草图的语义抽象效率仍具挑战。现有方法依赖参考图像、低层视觉特征或识别准确率,无法捕捉草图的核心特性——抽象性。为此,我们提出SEA(Sketch Evaluation metric for Abstraction efficiency),一种无参考的评估指标,衡量草图在保持语义可识别的前提下,对类别定义性视觉元素的经济表达程度。这些元素基于每类物体的常识知识生成。SEA利用视觉问答模型判断各元素是否存在,并输出反映语义保留程度的量化分数。为支持该指标,我们构建了CommonSketch,首个语义标注的草图数据集,包含300类共23,100张人工绘制草图,每张配有一段描述和元素级标注。实验表明,SEA与人类判断高度一致,能可靠区分抽象效率差异,CommonSketch则成为评估各类视觉-语言模型元素级理解能力的基准。

原文摘要 · Abstract (English)

A sketch is a distilled form of visual abstraction that conveys core concepts through simplified yet purposeful strokes while omitting extraneous detail. Despite its expressive power, quantifying the efficiency of semantic abstraction in sketches remains challenging. Existing evaluation methods that rely on reference images, low-level visual features, or recognition accuracy do not capture abstraction, the defining property of sketches. To address these limitations, we introduce SEA (Sketch Evaluation metric for Abstraction efficiency), a reference-free metric that assesses how economically a sketch represents class-defining visual elements while preserving semantic recognizability. These elements are derived per class from commonsense knowledge about features typically depicted in sketches. SEA leverages a visual question answering model to determine the presence of each element and returns a quantitative score that reflects semantic retention under visual economy. To support this metric, we present CommonSketch, the first semantically annotated sketch dataset, comprising 23,100 human-drawn sketches across 300 classes, each paired with a caption and element-level annotations. Experiments show that SEA aligns closely with human judgments and reliably discriminates levels of abstraction efficiency, while CommonSketch serves as a benchmark providing systematic evaluation of element-level sketch understanding across various vision-language models.

草图理解视觉问答抽象评估数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。