让视觉语言模型动态融合外部知识,提升问答与推理能力。
Dynamic Knowledge Integration for Enhanced Vision-Language Reasoning
- 用知识编码器+检索机制动态引入结构化与非结构化知识
- 在4个基准数据集上超越现有模型,人类评估更准确相关
- 适合需要外部知识的视觉问答、复杂推理场景
大型视觉语言模型(LVLMs)在多模态任务中表现优异,但其性能常受限于缺乏外部知识整合,难以处理知识密集型任务如视觉问答与推理。为此,我们提出一种新方法——自适应知识引导的大型视觉语言模型预训练(AKGP-LVLM),在预训练和微调阶段动态融入结构化与非结构化知识。该方法采用知识编码器表示外部知识,通过检索机制选择任务相关信息,并使用动态适配器有效对齐多模态与知识表征。我们在四个基准数据集上评估该方法,结果显著优于现有最先进模型。人工评估显示,模型输出具有更高的正确性与相关性。大量分析证实AKGP-LVLM具备鲁棒性、高效性与可扩展性,是解决现实世界知识密集型任务的有力方案。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multimodal tasks, but their performance is often constrained by the lack of external knowledge integration, limiting their ability to handle knowledge-intensive tasks such as visual question answering and reasoning. To address this challenge, we propose a novel method, Adaptive Knowledge-Guided Pretraining for Large Vision-Language Models (AKGP-LVLM), which dynamically incorporates structured and unstructured knowledge into LVLMs during pretraining and fine-tuning. Our approach employs a knowledge encoder to represent external knowledge, a retrieval mechanism to select task-relevant information, and a dynamic adaptor to align multimodal and knowledge representations effectively. We evaluate our method on four benchmark datasets, demonstrating significant performance improvements over state-of-the-art models. Furthermore, human evaluations highlight the superior correctness and relevance of our model's outputs. Extensive analyses confirm the robustness, efficiency, and scalability of AKGP-LVLM, making it a compelling solution for real-world knowledge-intensive tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。