通过模拟人类认知分解视觉推理,显著提升图表理解能力。
VisDoT : Enhancing Visual Reasoning through Human-Like Interpretation Grounding and Decomposition of Thought
- 将问题分解为感知与逻辑子问题,模仿人类思考过程。
- 在ChartQA上提升11.2%,在新基准VisDoTQA上提升33.2%。
- 适用于需要可解释视觉推理的复杂图表分析任务。
大型视觉语言模型在检测图表中的视觉基本元素并将其与语义表征对齐方面表现不佳,严重制约其在复杂视觉推理任务中的性能。这一感知基础缺失构成了图表推理的主要瓶颈。我们提出VisDoT框架,通过类人解释性感知锚定增强视觉推理。基于图形感知理论,形式化了包括位置与长度在内的四项感知任务。在此基础上,引入思维分解(DoT)提示策略,将问题逐层拆解为视觉感知子问题与逻辑推理子问题。使用VisDoT微调InternVL,在ChartQA上实现+11.2%的提升,并超越GPT-4o在更具挑战性的ChartQAPro基准上的表现。在新提出的VisDoTQA基准上,模型性能提升+33.2%。此外,在多种开放域VQA基准上均获得一致的零样本提升,验证了感知-逻辑分离策略在视觉问答中的通用性。VisDoT利用类人感知强化视觉锚定,实现了最先进的图表理解与可解释的视觉推理。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) struggle to reliably detect visual primitives in charts and align them with semantic representations, which severely limits their performance on complex visual reasoning. This lack of perceptual grounding constitutes a major bottleneck for chart-based reasoning. We propose VisDoT, a framework that enhances visual reasoning through human-like interpretation grounding. We formalize four perceptual tasks based on the theory of graphical perception, including position and length. Building on this foundation, we introduce Decomposition-of-Thought (DoT) prompting, which sequentially separates questions into visual perception sub-questions and logic sub-questions. Fine-tuning InternVL with VisDoT achieves a +11.2% improvement on ChartQA and surpasses GPT-4o on the more challenging ChartQAPro benchmark. On the newly introduced VisDoTQA benchmark, the model improves by +33.2%. Furthermore, consistent zero-shot gains on diverse open-domain VQA benchmarks confirm the generalizability of the perception-logic separation strategy for visual question answering. VisDoT leverages human-like perception to enhance visual grounding, achieving state-of-the-art chart understanding and interpretable visual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。