用图结构建模图表元素关系,提升问答准确率
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
- 构建视觉与文本双图结构,显式捕捉图表组件关系
- 通过图对比学习对齐多模态表示,显著提升模型性能
- 设计思维链提示减少幻觉,适合零样本场景应用
图表问答(ChartQA)面临图表元素异构性强、隐含数据模式细微等挑战。本文提出一种联合多模态场景图框架,显式建模图表组件及其底层结构之间的关系。该框架融合视觉图与文本图,捕获结构与语义特征,并采用图对比学习策略,在不同模态间对齐节点表示,使其作为软提示无缝融入Transformer解码器。此外,设计了一组定制化的思维链(Chain of Thought, CoT)提示,通过缓解幻觉现象,增强多模态大模型在零样本场景下的表现。在ChartQA、OpenCQA和ChartX等多个基准上的大量实验表明,所提方法显著提升性能,验证了其有效性。
原文摘要 · Abstract (English)
Chart question answering (ChartQA) is challenged by the heterogeneous composition of chart elements and the subtle data patterns they encode. This work introduces a novel joint multimodal scene graph framework that explicitly models the relationships among chart components and their underlying structures. The framework integrates both visual and textual graphs to capture structural and semantic characteristics, while a graph contrastive learning strategy aligns node representations across modalities enabling their seamless incorporation into a transformer decoder as soft prompts. Moreover, a set of tailored Chain of Thought (CoT) prompts is proposed to enhance multimodal large language models (MLLMs) in zero-s ot scenarios by mitigating hallucinations. Extensive evaluations on benchmarks including ChartQA, OpenCQA, and ChartX demonstrate significant performance improvements and validate the efficacy of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。