综述多模态大模型在视觉问答中的语言理解与推理进展
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
- 梳理图文理解与基于图像-问题信息的知识推理机制
- 总结视觉-语言预训练与多模态大模型的模态融合技术
- 适合关注多模态模型推理能力的研究者阅读
视觉问答(VQA)是融合自然语言处理与计算机视觉的挑战性任务,逐渐成为多模态大语言模型(MLLM)的基准测试任务。本综述旨在提供VQA发展的整体概览,并详细描述最新模型的高时效性进展。综述涵盖图像与文本的自然语言理解、基于图像-问题信息的知识推理模块,以及当前最先进的视觉-语言预训练模型和多模态大语言模型在模态信息提取与融合方面的最新成果。此外,系统回顾了知识推理在VQA中的进展,包括内部知识提取与外部知识引入。最后,列出主流VQA数据集与评估指标,并探讨未来研究方向。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) is a challenge task that combines natural language processing and computer vision techniques and gradually becomes a benchmark test task in multimodal large language models (MLLMs). The goal of our survey is to provide an overview of the development of VQA and a detailed description of the latest models with high timeliness. This survey gives an up-to-date synthesis of natural language understanding of images and text, as well as the knowledge reasoning module based on image-question information on the core VQA tasks. In addition, we elaborate on recent advances in extracting and fusing modal information with vision-language pretraining models and multimodal large language models in VQA. We also exhaustively review the progress of knowledge reasoning in VQA by detailing the extraction of internal knowledge and the introduction of external knowledge. Finally, we present the datasets of VQA and different evaluation metrics and discuss possible directions for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。