arXiv:2504.17547cs.CVcs.IR2025-04中稿 · TKDE, 20 pages, 5 …综述被引 7

系统梳理视觉问答中知识的生命周期,揭示模型如何用知识推理

A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task

  • 按知识表示、检索、推理三阶段分类现有方法
  • 指出跨模态对齐与噪声知识处理是核心挑战
  • 适合想了解知识增强视觉推理的研究者和应用开发者

基于知识的视觉问答(KB-VQA)在传统视觉问答基础上,不仅需理解图像和文本输入,还需整合广泛的知识,推动了多个实际应用的发展。其面临独特挑战:异构信息在多模态、多源数据间的对齐,从嘈杂或大规模知识库中检索相关知识,以及结合上下文进行复杂推理以得出答案。随着大语言模型(LLMs)的发展,KB-VQA系统发生显著演变,LLMs被用作强大的知识库、检索增强生成器和强推理引擎。尽管进展显著,但尚无系统性综述全面归纳现有方法。本综述旨在填补空白,建立KB-VQA方法的结构化分类体系,将系统划分为知识表示、知识检索和知识推理三大阶段。通过分析多种知识融合技术,识别持续存在的挑战,并提出有前景的未来研究方向,为推进KB-VQA模型及其应用奠定基础。

原文摘要 · Abstract (English)

Knowledge-based Vision Question Answering (KB-VQA) extends general Vision Question Answering (VQA) by not only requiring the understanding of visual and textual inputs but also extensive range of knowledge, enabling significant advancements across various real-world applications. KB-VQA introduces unique challenges, including the alignment of heterogeneous information from diverse modalities and sources, the retrieval of relevant knowledge from noisy or large-scale repositories, and the execution of complex reasoning to infer answers from the combined context. With the advancement of Large Language Models (LLMs), KB-VQA systems have also undergone a notable transformation, where LLMs serve as powerful knowledge repositories, retrieval-augmented generators and strong reasoners. Despite substantial progress, no comprehensive survey currently exists that systematically organizes and reviews the existing KB-VQA methods. This survey aims to fill this gap by establishing a structured taxonomy of KB-VQA approaches, and categorizing the systems into main stages: knowledge representation, knowledge retrieval, and knowledge reasoning. By exploring various knowledge integration techniques and identifying persistent challenges, this work also outlines promising future research directions, providing a foundation for advancing KB-VQA models and their applications.

视觉问答知识融合大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。