构建可分解视觉与推理能力的文档智能评测基准,助力模型精准定位短板。
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
- 将视觉复杂度与推理复杂度解耦,设计分级任务评估模型表现
- 覆盖19项文档任务、2.3千张图像,揭示模型在不同维度的优劣
- 适合研究文档理解、多模态模型优化的开发者和研究人员
多模态大语言模型的快速发展深刻影响了文档领域,催生了丰富的应用场景。然而现有评测基准难以准确识别模型弱点,也难以为系统性改进提供指引。为此,我们提出通用文档智能评测基准(GDI-Bench),包含9个关键场景、19项文档特定任务及2.3k张图像。通过解耦视觉复杂度与推理复杂度,构建分级任务体系,实现按难度评估性能,辅助模型弱点定位与优化指导。我们在GDI-Bench上评估多种开源与闭源模型,分别开展视觉与推理域的解耦分析,揭示其优势与不足。针对任务多样性,提出GDI-Model,采用保智训练策略缓解监督微调中的灾难性遗忘,强化基座模型固有短板。该模型在多个先前基准及GDI-Bench上均达领先水平。相关数据集与模型已开源于https://huggingface.co/GDIBench。
原文摘要 · Abstract (English)
The rapid advancement of multimodal large language models (MLLMs) has profoundly impacted the document domain, creating a wide array of application scenarios. This progress highlights the need for a comprehensive benchmark to evaluate these models' capabilities across various document-specific tasks. However, existing benchmarks often fail to locate specific model weaknesses or guide systematic improvements. To bridge this gap, we introduce a General Document Intelligence Benchmark (GDI-Bench), featuring 2.3k images across 9 key scenarios and 19 document-specific tasks. By decoupling visual complexity and reasoning complexity, the GDI-Bench structures graded tasks that allow performance assessment by difficulty, aiding in model weakness identification and optimization guidance. We evaluate various open-source and closed-source models on GDI-Bench, conducting decoupled analyses in the visual and reasoning domains, revealing their strengths and weaknesses. To address the diverse tasks and domains in the GDI-Bench, we propose a GDI-Model that mitigates catastrophic forgetting during the supervised fine-tuning (SFT) process through an intelligence-preserving training strategy, thereby reinforcing the inherent weaknesses of the base model. Our model achieves state-of-the-art performance on previous benchmarks and the GDI-Bench. Both our benchmark and models are or will be open-sourced on https://huggingface.co/GDIBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。