arXiv:2410.19144cs.CVcs.AI2024-10EMNLP被引 3

用视觉文字实体知识提升图文问答准确率

Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant

  • 引入视觉文字实体链接模块,结合图像上下文与文本信息进行精准匹配
  • 在三个数据集上平均比此前最优方法提升23.3%,刷新图文问答性能纪录
  • 适合研究多模态推理、知识增强模型的学者和开发者参考

我们重新审视基于文本的视觉问答(Text-KVQA),结合大型多模态模型(LMMs)的新进展,提出两项贡献:(i) 提出 VisTEL——一种基于原则的视觉文字实体链接方法。该模块利用先进的视觉文字识别引擎与大型多模态模型,联合利用图像中的上下文线索,将视觉文字实体与知识库实体准确关联。(ii) 提出 KaLMA——一个知识感知的大型多模态助手,通过为 LMM 增强图像中视觉文字实体的知识,实现更精确的回答。我们在多个基准上进行了全面实验,涵盖传统视觉问答、预大模型及当前主流 LMMs,并与以往最优方法对比。在 Text-KVQA 的三个数据集分叉上,我们的方法平均绝对提升 23.3%,确立新基准。代码已公开。

原文摘要 · Abstract (English)

We revisit knowledge-aware text-based visual question answering, also known as Text-KVQA, in the light of modern advancements in large multimodal models (LMMs), and make the following contributions: (i) We propose VisTEL - a principled approach to perform visual text entity linking. The proposed VisTEL module harnesses a state-of-the-art visual text recognition engine and the power of a large multimodal model to jointly reason using textual and visual context obtained using surrounding cues in the image to link the visual text entity to the correct knowledge base entity. (ii) We present KaLMA - a knowledge-aware large multimodal assistant that augments an LMM with knowledge associated with visual text entity in the image to arrive at an accurate answer. Further, we provide a comprehensive experimental analysis and comparison of our approach with traditional visual question answering, pre-large multimodal models, and large multimodal models, as well as prior top-performing approaches. Averaging over three splits of Text-KVQA, our proposed approach surpasses the previous best approach by a substantial 23.3% on an absolute scale and establishes a new state of the art. We make our implementation publicly available.

图文问答多模态知识增强视觉文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。