用真人看图轨迹优化大模型读图表,提升准确率与可解释性。
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
- 基于人类注视数据,引导模型聚焦图表关键区域。
- 在多模型上最高提升2.56个百分点准确率。
- 适合关注图表理解、模型可解释性的研究者。
图表是传递信息的重要视觉媒介。尽管大型视觉-语言模型(LVLMs)在图表问答(CQA)任务上取得进展,但模型常关注无关区域,导致理解偏差。本文提出ChartGaze,一个捕捉人类在图表推理任务中注视模式的新眼动数据集。通过对比人类与模型注意力,发现LVLMs常偏离人类注视路径,影响可解释性与准确率。为此,我们提出一种注视引导的注意力优化方法,使图像-文本注意力对齐人类注视点。该方法显著提升答案准确率与注意力一致性,在多个模型上最高达2.56个百分点的增益。结果表明,引入人类注视数据能有效提升图表导向型LVLM的推理质量与可解释性。
原文摘要 · Abstract (English)
Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particularly when models attend to irrelevant regions of the chart. In this work, we present ChartGaze, a new eye-tracking dataset that captures human gaze patterns during chart reasoning tasks. Through a systematic comparison of human and model attention, we find that LVLMs often diverge from human gaze, leading to reduced interpretability and accuracy. To address this, we propose a gaze-guided attention refinement that aligns image-text attention with human fixations. Our approach improves both answer accuracy and attention alignment, yielding gains of up to 2.56 percentage points across multiple models. These results demonstrate the promise of incorporating human gaze to enhance both the reasoning quality and interpretability of chart-focused LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。