梳理视觉问答领域十年演进,揭示关键模型与技术突破。
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
- 以注意力机制和视觉语言预训练为核心推动发展
- 强调Transformer架构在多模态理解中的核心作用
- 适合关注AI图像理解与跨模态研究的读者
视觉问答(VQA)是连接计算机视觉(CV)与自然语言处理(NLP)的交叉领域,使AI系统能够回答关于图像的问题。自2015年提出以来,随着深度学习、注意力机制及基于Transformer的模型发展,VQA迅速演进。本文回顾了VQA的发展历程,涵盖注意力机制、组合推理以及视觉-语言预训练方法的重大突破。文中重点介绍塑造VQA发展的关键模型、数据集与技术,突出变压器架构和多模态预训练在近期进展中的决定性作用。同时探讨了医疗等领域的专用应用,指出数据偏差、模型可解释性及常识推理等持续挑战。最后,展望大模型多模态趋势与外部知识融合,为未来方向提供洞见。本文旨在全面概述VQA的演进,呈现其现状与潜在进步。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) is an interdisciplinary field that bridges the gap between computer vision (CV) and natural language processing(NLP), enabling Artificial Intelligence(AI) systems to answer questions about images. Since its inception in 2015, VQA has rapidly evolved, driven by advances in deep learning, attention mechanisms, and transformer-based models. This survey traces the journey of VQA from its early days, through major breakthroughs, such as attention mechanisms, compositional reasoning, and the rise of vision-language pre-training methods. We highlight key models, datasets, and techniques that shaped the development of VQA systems, emphasizing the pivotal role of transformer architectures and multimodal pre-training in driving recent progress. Additionally, we explore specialized applications of VQA in domains like healthcare and discuss ongoing challenges, such as dataset bias, model interpretability, and the need for common-sense reasoning. Lastly, we discuss the emerging trends in large multimodal language models and the integration of external knowledge, offering insights into the future directions of VQA. This paper aims to provide a comprehensive overview of the evolution of VQA, highlighting both its current state and potential advancements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。