综述视觉语言模型在图表理解中的最新进展与未来方向
Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends
- 基于Transformer的端到端框架显著提升图表理解性能
- 32篇最新研究被筛选分析,涵盖多任务与单任务解决方案
- 适合关注图表智能、多模态模型与可解释AI的研究者
近年来,视觉-语言任务尤其是图表交互受到广泛关注。这些任务具有多模态特性,需同时处理图表图像、文本说明、底层数据表及用户查询。传统图表理解(CU)依赖启发式规则,而近年来融合Transformer架构的方法显著提升了性能。本文系统回顾了2020年1月至2024年6月间32项代表性研究,遵循PRISMA指南,聚焦采用Transformer的端到端(E2E)框架。按认知任务分为三层范式,框架分为单任务与多任务两类,其中多任务框架进一步探索预训练与提示工程方法。综述了主流架构、数据集与预训练任务。尽管进展显著,仍面临OCR依赖、低分辨率图像处理及视觉推理能力不足等挑战。未来方向包括构建更鲁棒的基准、优化模型效率,以及融合可解释AI与真实/合成数据的平衡策略。
原文摘要 · Abstract (English)
In recent years, interest in vision-language tasks has grown, especially those involving chart interactions. These tasks are inherently multimodal, requiring models to process chart images, accompanying text, underlying data tables, and often user queries. Traditionally, Chart Understanding (CU) relied on heuristics and rule-based systems. However, recent advancements that have integrated transformer architectures significantly improved performance. This paper reviews prominent research in CU, focusing on State-of-The-Art (SoTA) frameworks that employ transformers within End-to-End (E2E) solutions. Relevant benchmarking datasets and evaluation techniques are analyzed. Additionally, this article identifies key challenges and outlines promising future directions for advancing CU solutions. Following the PRISMA guidelines, a comprehensive literature search is conducted across Google Scholar, focusing on publications from Jan'20 to Jun'24. After rigorous screening and quality assessment, 32 studies are selected for in-depth analysis. The CU tasks are categorized into a three-layered paradigm based on the cognitive task required. Recent advancements in the frameworks addressing various CU tasks are also reviewed. Frameworks are categorized into single-task or multi-task based on the number of tasks solvable by the E2E solution. Within multi-task frameworks, pre-trained and prompt-engineering-based techniques are explored. This review overviews leading architectures, datasets, and pre-training tasks. Despite significant progress, challenges remain in OCR dependency, handling low-resolution images, and enhancing visual reasoning. Future directions include addressing these challenges, developing robust benchmarks, and optimizing model efficiency. Additionally, integrating explainable AI techniques and exploring the balance between real and synthetic data are crucial for advancing CU research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。