系统梳理视觉问答发展脉络,涵盖方法、数据与应用。
Visual question answering: from early developments to recent advances -- a survey
- 按架构设计分类,构建VQA研究体系
- 总结深度学习与大视觉语言模型进展
- 适合入门者与跨领域研究者参考
视觉问答(VQA)是一项融合图像与语言处理的多模态研究领域,旨在让机器通过特征提取、目标检测、文本嵌入、自然语言理解与语言生成等技术回答图像相关问题。随着多模态数据研究的发展,VQA因其在交互式教育工具、医学影像诊断、客户服务、娱乐及社交媒体图文生成等场景中的广泛应用而备受关注,同时在辅助视障人士生成图像描述方面具有重要价值。本文提出VQA架构的分类体系,基于设计选择与核心组件进行归纳,便于比较分析。综述主流深度学习方法,探讨大视觉语言模型(LVLMs)在多模态任务如VQA中的成功应用。文章还梳理了常用数据集与评估指标,分析实际应用场景,并指出当前挑战与未来研究方向,为研究人员与实践者提供全面参考。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) is an evolving research field aimed at enabling machines to answer questions about visual content by integrating image and language processing techniques such as feature extraction, object detection, text embedding, natural language understanding, and language generation. With the growth of multimodal data research, VQA has gained significant attention due to its broad applications, including interactive educational tools, medical image diagnosis, customer service, entertainment, and social media captioning. Additionally, VQA plays a vital role in assisting visually impaired individuals by generating descriptive content from images. This survey introduces a taxonomy of VQA architectures, categorizing them based on design choices and key components to facilitate comparative analysis and evaluation. We review major VQA approaches, focusing on deep learning-based methods, and explore the emerging field of Large Visual Language Models (LVLMs) that have demonstrated success in multimodal tasks like VQA. The paper further examines available datasets and evaluation metrics essential for measuring VQA system performance, followed by an exploration of real-world VQA applications. Finally, we highlight ongoing challenges and future directions in VQA research, presenting open questions and potential areas for further development. This survey serves as a comprehensive resource for researchers and practitioners interested in the latest advancements and future
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。