arXiv:2602.10138cs.CVcs.AI2026-02综述被引 1

系统梳理多模态大模型在图表理解中的演进与瓶颈

Multimodal Information Fusion for Chart Understanding: A Survey of MLLMs -- Evolution, Limitations, and Cognitive Enhancement

  • 构建图表图文融合的分类体系与基准测试框架
  • 揭示现有模型在感知与推理上的双重缺陷
  • 适合研究多模态认知增强与智能分析的学者

图表理解是典型的多模态信息融合任务,需无缝整合图形与文本数据以提取语义。多模态大语言模型(MLLMs)的出现彻底改变了该领域,但基于MLLM的图表分析仍缺乏系统性组织。本综述通过梳理核心组件,为这一新兴前沿提供全面路线图。首先分析图表中视觉与语言信息融合的基本挑战;其次对下游任务与数据集进行分类,提出原创的规范性与非规范性基准分类法,凸显领域扩展趋势;随后系统回顾方法论演化,从经典深度学习技术追溯至当前最先进的MLLM范式及其复杂融合策略;最后批判性审视现有模型的局限,尤其在感知与推理能力方面的不足,指出未来方向,包括高级对齐技术与强化学习驱动的认知增强。本综述旨在帮助研究者与实践者建立对MLLM如何重塑图表信息融合的结构化认知,并推动更鲁棒可靠的系统发展。

原文摘要 · Abstract (English)

Chart understanding is a quintessential information fusion task, requiring the seamless integration of graphical and textual data to extract meaning. The advent of Multimodal Large Language Models (MLLMs) has revolutionized this domain, yet the landscape of MLLM-based chart analysis remains fragmented and lacks systematic organization. This survey provides a comprehensive roadmap of this nascent frontier by structuring the domain's core components. We begin by analyzing the fundamental challenges of fusing visual and linguistic information in charts. We then categorize downstream tasks and datasets, introducing a novel taxonomy of canonical and non-canonical benchmarks to highlight the field's expanding scope. Subsequently, we present a comprehensive evolution of methodologies, tracing the progression from classic deep learning techniques to state-of-the-art MLLM paradigms that leverage sophisticated fusion strategies. By critically examining the limitations of current models, particularly their perceptual and reasoning deficits, we identify promising future directions, including advanced alignment techniques and reinforcement learning for cognitive enhancement. This survey aims to equip researchers and practitioners with a structured understanding of how MLLMs are transforming chart information fusion and to catalyze progress toward more robust and reliable systems.

图表理解多模态大模型认知增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。