arXiv:2605.01882cs.CVcs.AI2026-05

针对密集图表的细粒度推理难题,提出视觉聚焦新模型

Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts

论文配图:Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts
图 1 · 摘自论文原文
  • 用视觉聚焦思维链显式关联推理与关键视觉线索
  • 在多个基准上超越现有模型,尤其在高信息密度图表上提升显著
  • 适合需要精准分析复杂图表的研究者和开发者

多模态大语言模型在图表理解任务中展现潜力,但在包含多个子图、图例和密集标注的高信息密度(HID)图表上仍面临三大挑战:(1)细粒度感知能力有限导致关键视觉线索遗漏;(2)冗余或噪声信息影响多模态推理性能;(3)缺乏与视觉信息量相适应的深度推理机制。为此,我们提出新型视觉聚焦细粒度图表推理模型 Chart-FR1,以提升感知、聚焦效率与自适应深度推理能力。具体地,提出 Focus-CoT——一种将推理步骤显式关联至局部图像区域和 OCR 信号等关键视觉线索的视觉聚焦思维链;在此基础上,引入 Focus-GRPO,一种基于信息效率奖励的聚焦强化学习算法,通过压缩冗余视觉信息实现高效聚焦,并设计自适应 KL 惩罚机制,随发现更多视觉线索灵活调控推理深度。此外,为填补 HID 图表评估基准空白,构建了 HID-Chart 基准,其信息密度指标可有效评估细粒度推理能力。在多个图表基准上的大量实验表明,Chart-FR1 在图表理解与推理任务中优于当前最优 MLLMs。代码已开源:https://github.com/phkhub/Chart-FR1。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have shown considerable potential in chart understanding and reasoning tasks. However, they still struggle with high information density (HID) charts characterized by multiple subplots, legends, and dense annotations due to three major challenges: (1) limited fine-grained perception results in the omission of critical visual cues; (2) redundant or noisy visual information undermines the performance of multimodal reasoning; (3) lack of adaptive deep reasoning relative to the amount of visual information. To tackle these challenges, we present a novel focus-driven fine-grained chart reasoning model, Chart-FR1, to improve perception, focusing efficiency, and adaptive deep reasoning on HID charts. Specifically, we propose Focus-CoT, a visual focusing chain-of-thought that enhances fine-grained perception by explicitly linking reasoning steps to key visual cues, such as local image regions and OCR signals. Building on this, we introduce Focus-GRPO, a focus-driven reinforcement learning algorithm with an information-efficiency reward that compresses redundant visual information for efficient focusing, and an adaptive KL penalty mechanism that enables flexible control over reasoning depth as more visual cues are discovered. Furthermore, to fill the gap in benchmarks for HID charts, we build HID-Chart, a challenging benchmark with an information-density metric designed to evaluate fine-grained chart reasoning capabilities. Extensive experiments on multiple chart benchmarks demonstrate that Chart-FR1 outperforms state-of-the-art MLLMs in chart understanding and reasoning. Code is available at https://github.com/phkhub/Chart-FR1.

图表理解细粒度推理视觉聚焦多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。