用视觉属性引导众包标注,减少主观偏差。
Crowdsourcing of Real-world Image Annotation via Visual Properties

- 基于物体类别层级与反馈动态提问
- 显著降低标注主观性,提升一致性
- 适合需要高精度标注的视觉任务
数据驱动的人工智能发展暴露出目标识别数据集的固有局限。主要问题源于语义鸿沟,导致视觉数据与语言描述之间存在复杂的多对多映射关系,从而影响计算机视觉性能。本文提出一种融合知识表示、自然语言处理与计算机视觉的图像标注方法,通过施加视觉属性约束来减少标注者主观性。设计了一种交互式众包框架,根据预定义的物体类别层次结构和标注者反馈动态生成问题,以视觉属性为导向引导图像标注。实验验证了该方法的有效性,同时分析了标注者反馈以优化众包设置。
原文摘要 · Abstract (English)
Recent advances in data-centric artificial intelligence highlight inherent limitations in object recognition datasets. One of the primary issues stems from the semantic gap problem, which results in complex many-to-many mappings between visual data and linguistic descriptions. This bias adversely affects performance in computer vision tasks. This paper proposes an image annotation methodology that integrates knowledge representation, natural language processing, and computer vision techniques, aiming to reduce annotator subjectivity by applying visual property constraints. We introduce an interactive crowdsourcing framework that dynamically asks questions based on a predefined object category hierarchy and annotator feedback, guiding image annotation by visual properties. Experiments demonstrate the effectiveness of this methodology, and annotator feedback is discussed to optimize the crowdsourcing setup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。