综述图文理解多模态大模型发展,梳理架构与前沿进展。
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review
- 系统梳理文本丰富图像理解模型的演进时间线与架构设计
- 对比主流基准上多个模型的性能表现
- 指出未来方向与当前技术瓶颈,适合研究者快速入门
多模态大语言模型(MLLMs)的兴起为文本丰富图像理解(TIU)领域带来了新维度,模型展现出令人印象深刻且鼓舞人心的性能。然而,其快速演进和广泛应用使得追踪最新进展变得愈发困难。为此,我们提出一项系统且全面的综述,以促进对TIU MLLMs的进一步研究。首先,我们梳理了几乎所有TIU MLLMs的时间线、架构与处理流程;其次,回顾了选定模型在主流基准上的表现;最后,探讨了该领域的潜在方向、挑战与局限性。
原文摘要 · Abstract (English)
The recent emergence of Multi-modal Large Language Models (MLLMs) has introduced a new dimension to the Text-rich Image Understanding (TIU) field, with models demonstrating impressive and inspiring performance. However, their rapid evolution and widespread adoption have made it increasingly challenging to keep up with the latest advancements. To address this, we present a systematic and comprehensive survey to facilitate further research on TIU MLLMs. Initially, we outline the timeline, architecture, and pipeline of nearly all TIU MLLMs. Then, we review the performance of selected models on mainstream benchmarks. Finally, we explore promising directions, challenges, and limitations within the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。