评测并提升大模型对视觉图结构的理解与推理能力
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
- 构建22项任务的基准VGCure,系统评估模型图理解能力
- 14个大模型在复杂关系任务上表现普遍较差,平均准确率不足50%
- 提出结构感知微调框架,显著提升模型对图结构的鲁棒性
大型视觉语言模型(LVLMs)在多种任务中表现出色。然而,近期研究发现,当处理视觉图时,这些模型存在显著局限。为探究其原因,我们提出了VGCure——一个涵盖22项任务的综合性基准,用于评估LVLMs在基础图理解与推理方面的能力。对14个主流LVLMs的广泛评估显示,它们在涉及关系或结构复杂信息的任务中表现薄弱。基于此发现,我们设计了一种结构感知微调框架,通过三项自监督学习任务赋予模型结构学习能力。实验表明,该方法有效提升了模型在基础及下游图学习任务中的表现,并增强了其对复杂视觉图的鲁棒性。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across diverse tasks. Despite great success, recent studies show that LVLMs encounter substantial limitations when engaging with visual graphs. To study the reason behind these limitations, we propose VGCure, a comprehensive benchmark covering 22 tasks for examining the fundamental graph understanding and reasoning capacities of LVLMs. Extensive evaluations conducted on 14 LVLMs reveal that LVLMs are weak in basic graph understanding and reasoning tasks, particularly those concerning relational or structurally complex information. Based on this observation, we propose a structure-aware fine-tuning framework to enhance LVLMs with structure learning abilities through three self-supervised learning tasks. Experiments validate the effectiveness of our method in improving LVLMs' performance on fundamental and downstream graph learning tasks, as well as enhancing their robustness against complex visual graphs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。