用可定制图表生成器测试大视觉语言模型在图形分析中的表现差异。
VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models
- 构建可视化图变异基准生成器,支持七类图任务评估。
- 990张图测试显示,节点重叠等视觉缺陷显著降低模型性能。
- 适合研究视觉语言模型在复杂图形任务中的鲁棒性与局限性。
大型视觉语言模型(LVLMs)在处理抽象视觉任务方面展现出巨大潜力。几何结构,尤其是具有内在灵活性和复杂性的图,是评估这些模型预测能力的理想基准。尽管人类观察者能轻易识别细微视觉特征并进行准确分析,我们的研究发现,当前最先进的LVLMs在特定视觉图场景中存在持续性局限,尤其在面对风格变化时。为此,我们提出VisGraphVar(视觉图变异),一个可定制的基准生成器,能够生成涵盖七类任务(检测、分类、分割、模式识别、链接预测、推理、匹配)的图图像,系统评估单个LVLM的优势与不足。我们利用VisGraphVar生成990张图,并采用零样本与思维链两种提示策略,评估六种LVLM。结果表明,图像中视觉属性的变化(如节点标签和布局)以及故意引入的视觉瑕疵(如节点重叠)显著影响模型表现。本研究强调了在图相关任务上进行全面评估的重要性,超越单一推理能力。VisGraphVar为开发更可靠、更鲁棒的高级视觉图分析系统提供了宝贵见解。
原文摘要 · Abstract (English)
The fast advancement of Large Vision-Language Models (LVLMs) has shown immense potential. These models are increasingly capable of tackling abstract visual tasks. Geometric structures, particularly graphs with their inherent flexibility and complexity, serve as an excellent benchmark for evaluating these models' predictive capabilities. While human observers can readily identify subtle visual details and perform accurate analyses, our investigation reveals that state-of-the-art LVLMs exhibit consistent limitations in specific visual graph scenarios, especially when confronted with stylistic variations. In response to these challenges, we introduce VisGraphVar (Visual Graph Variability), a customizable benchmark generator able to produce graph images for seven distinct task categories (detection, classification, segmentation, pattern recognition, link prediction, reasoning, matching), designed to systematically evaluate the strengths and limitations of individual LVLMs. We use VisGraphVar to produce 990 graph images and evaluate six LVLMs, employing two distinct prompting strategies, namely zero-shot and chain-of-thought. The findings demonstrate that variations in visual attributes of images (e.g., node labeling and layout) and the deliberate inclusion of visual imperfections, such as overlapping nodes, significantly affect model performance. This research emphasizes the importance of a comprehensive evaluation across graph-related tasks, extending beyond reasoning alone. VisGraphVar offers valuable insights to guide the development of more reliable and robust systems capable of performing advanced visual graph analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。