用外部符号系统增强视觉语言模型,提升推理与可解释性。
Augmented Vision-Language Models: A Systematic Review
- 将预训练视觉语言模型与外部符号系统结合,实现神经符号融合。
- 增强模型对新信息的适应能力,减少重训练需求。
- 适合需要可解释性与逻辑推理的AI应用研究者。
近年来,视觉-语言机器学习模型在大规模非结构化数据上训练后,展现出强大的自然语言理解与视觉场景感知能力。然而,这种训练范式难以提供输出的可解释性,需重新训练才能融入新信息,资源消耗高,且在逻辑推理方面表现不足。一种有前景的解决方案是将神经网络与外部符号信息系统结合,构建神经符号系统,从而增强推理与记忆能力。这类系统能提供更可解释的输出,并可在不进行大量重训练的情况下整合新知识。以强大的预训练视觉-语言模型(VLMs)为核心,辅以外部系统,是一种实现神经符号集成的务实路径。本文系统综述旨在分类梳理通过与外部符号信息系统交互来提升视觉-语言理解的技术方法。
原文摘要 · Abstract (English)
Recent advances in visual-language machine learning models have demonstrated exceptional ability to use natural language and understand visual scenes by training on large, unstructured datasets. However, this training paradigm cannot produce interpretable explanations for its outputs, requires retraining to integrate new information, is highly resource-intensive, and struggles with certain forms of logical reasoning. One promising solution involves integrating neural networks with external symbolic information systems, forming neural symbolic systems that can enhance reasoning and memory abilities. These neural symbolic systems provide more interpretable explanations to their outputs and the capacity to assimilate new information without extensive retraining. Utilizing powerful pre-trained Vision-Language Models (VLMs) as the core neural component, augmented by external systems, offers a pragmatic approach to realizing the benefits of neural-symbolic integration. This systematic literature review aims to categorize techniques through which visual-language understanding can be improved by interacting with external symbolic information systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。