arXiv:2602.23589cs.CVcs.AI2026-02

用合成图增强视觉语言模型对流程图的细微结构理解能力

Pseudo Contrastive Learning for Diagram Comprehension in Multimodal Models

  • 用随机文本元素生成伪对比样本,模拟图示结构差异
  • 在流程图匹配和问答任务上显著优于标准CLIP和硬负样本CLIP
  • 适合需要精准理解技术图表的科研与工程场景

近期的多模态模型(如对比语言-图像预训练,CLIP)在对齐视觉与语言表征方面表现出色。然而,在微小视觉差异蕴含重要语义的领域(如图示理解),由于模型对细粒度结构变化敏感性不足,仍面临挑战。本文提出一种新型训练范式,旨在提升视觉-语言模型在图示理解上的能力。方法通过图示渲染器生成伪对比样本,利用随机选取的文本元素构造合成图示,突出图示中结构差异,无需修改原始数据。将这些伪对比样本融入训练目标后,模型能更精确地捕捉语义一致的图示结构。在流程图基准数据集上的实证评估显示,该方法在图像-文本匹配与视觉问答任务中均显著优于标准CLIP及硬负样本CLIP训练方式。结果表明,领域特定的训练策略对推动图示理解具有重要价值。

原文摘要 · Abstract (English)

Recent multimodal models such as Contrastive Language-Image Pre-training (CLIP) have shown remarkable ability to align visual and linguistic representations. However, domains where small visual differences carry large semantic significance, such as diagram understanding, remain challenging due to the models' limited sensitivity to fine-grained structural variations. We propose a new training paradigm designed to enhance diagram comprehension in vision-language models. Our approach introduces pseudo contrastive samples generated by a diagram renderer that creates synthetic diagrams using randomly picked text elements. These samples highlight structural differences in diagrammatic imagery without requiring any modification or editing of the original data. By incorporating these pseudo contrastive samples into the training objective, the model learns to capture more precise and semantically consistent diagram structures. Empirical evaluations on a benchmark dataset of flowcharts demonstrate substantial improvements over standard CLIP and hard-negative CLIP training in both image-text matching and visual question answering tasks. The results underscore the value of domain-specific training strategies and contribute to advancing diagrammatic understanding within the broader context of vision-language learning.

图示理解对比学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。