arXiv:2412.06720cs.CVcs.CL2024-12被引 3

用图像区域引导实体链接,提升图文匹配准确率。

VP-MEL: Visual Prompts Guided Multimodal Entity Linking

  • 以图像区域为提示,融合图文信息进行实体链接
  • 在新数据集VPWiki上,模型性能超越基线方法
  • 适合研究视觉提示与多模态知识对齐的学者

多模态实体链接(MEL)旨在将图文上下文中的提及项关联到知识库中的对应实体,近年来应用广泛。然而,现有方法主要依赖文本提及词作为检索线索,难以有效利用图像与文本的联合信息,导致在关注图像对象或文本缺失提及词时链接不准。为此,本文提出视觉提示引导的多模态实体链接(VP-MEL)任务:给定图文对,需将图像中标记区域(即视觉提示)链接至知识库中对应的实体。为此,我们构建了专用数据集VPWiki,并提出IIER框架,通过视觉提示增强视觉特征提取,结合预训练Detective-VLM模型捕捉潜在语义信息。在VPWiki上的实验表明,IIER在多个基准上均优于基线方法。

原文摘要 · Abstract (English)

Multimodal entity linking (MEL), a task aimed at linking mentions within multimodal contexts to their corresponding entities in a knowledge base (KB), has attracted much attention due to its wide applications in recent years. However, existing MEL methods often rely on mention words as retrieval cues, which limits their ability to effectively utilize information from both images and text. This reliance causes MEL to struggle with accurately retrieving entities in certain scenarios, especially when the focus is on image objects or mention words are missing from the text. To solve these issues, we introduce a Visual Prompts guided Multimodal Entity Linking (VP-MEL) task. Given a text-image pair, VP-MEL aims to link a marked region (i.e., visual prompt) in an image to its corresponding entities in the knowledge base. To facilitate this task, we present a new dataset, VPWiki, specifically designed for VP-MEL. Furthermore, we propose a framework named IIER, which enhances visual feature extraction using visual prompts and leverages the pretrained Detective-VLM model to capture latent information. Experimental results on the VPWiki dataset demonstrate that IIER outperforms baseline methods across multiple benchmarks for the VP-MEL task.

多模态实体链接视觉提示知识库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。