arXiv:2505.19076cs.CV2025-05NeurIPS被引 12

让AI像人一样在图表上画画推理,边画边改,理解更准。

ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding

  • 用程序化绘图在图表上标注推理步骤,实现视觉化思考
  • 在多个基准上表现优于现有模型,尤其在复杂图表任务中
  • 适合需要可解释、交互式图表分析的科研与商业场景

图表是承载复杂数据的高密度可视化媒介,是信息提取与分析的关键工具。现有多模态大模型在自动图表理解方面面临挑战,因其需精准复杂的视觉推理能力。当前基于逐步推理的模型主要依赖文本逻辑,难以修正因视觉理解错误导致的偏差,缺乏利用多模态交互深化理解的能力。受人类认知行为启发,我们提出 ChartSketcher,一种基于多模态反馈的逐步推理方法。该模型采用 Sketch-CoT 机制,使多模态大语言模型能通过程序化绘图库将中间推理步骤直接标注于图表之上,并迭代地将这些视觉注释反馈至推理过程,实现推理的视觉锚定与逐步优化。我们采用两阶段训练策略:先进行冷启动阶段以学习基于绘图的推理模式,再通过离策略强化学习提升反思能力与泛化性能。实验表明,ChartSketcher 在图表理解基准和通用视觉任务上均取得良好表现,提供了一种交互性强、可解释的图表理解新范式。

原文摘要 · Abstract (English)

Charts are high-density visualization carriers for complex data, serving as a crucial medium for information extraction and analysis. Automated chart understanding poses significant challenges to existing multimodal large language models (MLLMs) due to the need for precise and complex visual reasoning. Current step-by-step reasoning models primarily focus on text-based logical reasoning for chart understanding. However, they struggle to refine or correct their reasoning when errors stem from flawed visual understanding, as they lack the ability to leverage multimodal interaction for deeper comprehension. Inspired by human cognitive behavior, we propose ChartSketcher, a multimodal feedback-driven step-by-step reasoning method designed to address these limitations. ChartSketcher is a chart understanding model that employs Sketch-CoT, enabling MLLMs to annotate intermediate reasoning steps directly onto charts using a programmatic sketching library, iteratively feeding these visual annotations back into the reasoning process. This mechanism enables the model to visually ground its reasoning and refine its understanding over multiple steps. We employ a two-stage training strategy: a cold start phase to learn sketch-based reasoning patterns, followed by off-policy reinforcement learning to enhance reflection and generalization. Experiments demonstrate that ChartSketcher achieves promising performance on chart understanding benchmarks and general vision tasks, providing an interactive and interpretable approach to chart comprehension.

图表理解多模态推理视觉标注交互式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。