用图书馆学分类法改进视觉数据标注,缩小图像与语义的差距
What can Computer Vision learn from Ranganathan?
- 借鉴拉冈纳坦分类原则重构图像标注体系
- vTelos方法使标注准确率显著提升
- 适合数据集设计者和标注流程优化者参考
计算机视觉中的语义鸿沟问题源于视觉与语言语义的不匹配,导致数据集设计和评估基准存在缺陷。本文提出,可借鉴S.R. 拉冈纳坦的分类原则,为解决该问题提供系统性起点,并指导高质量视觉数据集的设计。文中阐明这些原则经适当调整后,构成了vTelos视觉标注方法的核心理念。此外,论文还简要呈现了实验结果,显示vTelos在标注质量与模型准确率上均有提升,验证了其有效性。
原文摘要 · Abstract (English)
The Semantic Gap Problem (SGP) in Computer Vision (CV) arises from the misalignment between visual and lexical semantics leading to flawed CV dataset design and CV benchmarks. This paper proposes that classification principles of S.R. Ranganathan can offer a principled starting point to address SGP and design high-quality CV datasets. We elucidate how these principles, suitably adapted, underpin the vTelos CV annotation methodology. The paper also briefly presents experimental evidence showing improvements in CV annotation and accuracy, thereby, validating vTelos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。