arXiv:2510.13675cs.CVcs.LG2025-10AAAI被引 4

用知识图谱增强视觉实体识别,让模型看图识物更准、更懂未知概念。

Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning

  • 通过对比学习将图像与文本映射到知识图谱的语义空间中
  • 在未见实体上准确率提升10.5%,小模型性能超越大模型
  • 适合需要零样本识别和长尾分布处理的研究者

开放域视觉实体识别旨在将图像中出现的实体链接到如Wikidata般庞大且动态演化的现实世界概念集合。与标签集固定的分类任务不同,该任务在开放集条件下运行,训练时多数目标实体未见过,且服从长尾分布,导致监督有限、视觉模糊、语义歧义等问题。本文提出基于知识引导的对比学习框架(KnowCoL),将图像与文本描述统一映射至由Wikidata结构化信息支撑的共享语义空间。通过抽象至概念层面,模型利用实体描述、类型层次与关系上下文实现零样本识别。在包含Wikidata ID作为标签空间的大规模开放域视觉识别数据集OVEN上评估,实验表明结合视觉、文本与结构化知识显著提升精度,尤其对罕见及未见实体效果突出。最小模型在未见实体上准确率比当前最优方法提升10.5%,尽管体积仅为后者的1/35。

原文摘要 · Abstract (English)

Open-domain visual entity recognition aims to identify and link entities depicted in images to a vast and evolving set of real-world concepts, such as those found in Wikidata. Unlike conventional classification tasks with fixed label sets, it operates under open-set conditions, where most target entities are unseen during training and exhibit long-tail distributions. This makes the task inherently challenging due to limited supervision, high visual ambiguity, and the need for semantic disambiguation. We propose a Knowledge-guided Contrastive Learning (KnowCoL) framework that combines both images and text descriptions into a shared semantic space grounded by structured information from Wikidata. By abstracting visual and textual inputs to a conceptual level, the model leverages entity descriptions, type hierarchies, and relational context to support zero-shot entity recognition. We evaluate our approach on the OVEN benchmark, a large-scale open-domain visual recognition dataset with Wikidata IDs as the label space. Our experiments show that using visual, textual, and structured knowledge greatly improves accuracy, especially for rare and unseen entities. Our smallest model improves the accuracy on unseen entities by 10.5% compared to the state-of-the-art, despite being 35 times smaller.

视觉识别知识图谱零样本对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。