针对图表理解优化视觉语言模型,提升结构化信息识别能力。
Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models
- 用结构感知的对比学习,引入两类专为图表设计的损失函数。
- 在流程图数据集上,图文匹配与视觉问答任务性能显著优于CLIP。
- 适合需要精准理解技术图表、流程图等结构化图像的研究者。
视觉语言模型(如CLIP)在对齐视觉与语言表示方面表现卓越,但在处理流程图等结构化符号图像时存在局限。本文提出一种新型训练范式,通过引入“难样本”和两种专为图表结构设计的对比损失函数,增强模型对图表内容的结构化与语义一致性理解。该方法在流程图基准数据集上进行了验证,结果表明其在图文匹配和视觉问答任务中均显著优于标准CLIP及传统硬负样本对比学习方法。研究强调了针对特定任务设计训练策略的重要性,推动了视觉语言模型在图表理解领域的进展。
原文摘要 · Abstract (English)
Multimodal models, such as the Contrastive Language-Image Pre-training (CLIP) model, have demonstrated remarkable success in aligning visual and linguistic representations. However, these models exhibit limitations when applied to specialised visual domains, such as diagrams, which encode structured, symbolic information distinct from that of natural imagery. In this paper, we introduce a novel training paradigm explicitly designed to enhance the comprehension of diagrammatic images within vision-language models. Our approach uses ``hard'' samples for our proposed contrastive learning that incorporates two specialised loss functions that leverage the inherent structural properties of diagrams. By integrating these objectives into model training, our method enables models to develop a more structured and semantically coherent understanding of diagrammatic content. We empirically validate our approach on a benchmark dataset of flowcharts, as a representative class of diagrammatic imagery, demonstrating substantial improvements over standard CLIP and conventional hard negative CLIP learning paradigms for both image-text matching and visual question answering tasks. Our findings underscore the significance of tailored training strategies for specialised tasks and contribute to advancing diagrammatic understanding within the broader landscape of vision-language integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。