arXiv:2604.16517cs.CVcs.CL2026-04

小模型加知识图谱,让视觉语言模型更准更省

SmoGVLM: A Small, Graph-enhanced Vision-Language Model

论文配图:SmoGVLM: A Small, Graph-enhanced Vision-Language Model
图 1 · 摘自论文原文
  • 用图神经网络融合知识图谱与图文信息
  • 1.3B小模型性能提升16.24%,超越13B大模型
  • 适合资源有限但需精准推理的场景

大型视觉语言模型(VLM)在多模态任务中表现优异,但常出现幻觉且知识推理缺乏准确锚定。我们提出SmoGVLM,一种小型、图增强的视觉语言模型,通过图神经网络将结构化知识与视觉和文本模态融合。我们在从极小(1.3B)到大型(13B)的不同模型规模上验证了该方法。结果表明,采用本方法训练后,小型模型性能最高提升16.24%,甚至超过更大规模的VLM及强基线微调模型。这些发现凸显了结构化知识增强在高效、小型多模态推理系统中的潜力。

原文摘要 · Abstract (English)

Large vision-language models (VLMs) achieve strong performance on multimodal tasks but often suffer from hallucination and poor grounding in knowledge-intensive reasoning. We propose SmoGVLM, a small, graph-enhanced VLM that integrates structured knowledge with visual and textual modalities, using Graph Neural Networks. We investigate the effects of our method across a range of model sizes, from tiny (1.3B) to large (13B) models. Our results demonstrate that, when trained using our approach, a small model can achieve performance gains upto 16.24%, and surpass its larger counterparts, outperforming larger VLMs and strong fine-tuned baselines. These results highlight the potential of structured knowledge augmentation for efficient, smaller-scale multimodal reasoning systems.

视觉语言模型知识图谱小模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。