构建首个中英双语复杂文档理解大模型基准
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
- 用大模型自动生成多代理文档数据流水线
- 覆盖7.2万张图像与60万+问答对,支持多任务标注
- 适合研究真实场景文档理解的开发者和研究员
尽管视觉语言模型进展迅速,现有评测基准仍以英文为主、布局简单、任务有限,难以评估复杂版式文档的理解能力。为此,我们提出DocWeaver多智能体生成管道,利用大模型自动构建新基准。最终产出MosaicDoc——一个大规模、中英双语的视觉丰富文档理解基准。数据源自报纸与杂志,涵盖多栏、非曼哈顿布局,来自196家出版商的丰富风格,包含OCR、VQA、阅读顺序与定位等多任务标注。该数据集含7.2万张图像和超过60万组问答对,为领域提供权威评测标准。我们在该基准上对主流模型进行评估,揭示其在真实文档复杂性面前的局限性,并为未来研究指明方向。
原文摘要 · Abstract (English)
Despite the rapid progress of Vision-Language Models (VLMs), their capabilities are inadequately assessed by existing benchmarks, which are predominantly English-centric, feature simplistic layouts, and support limited tasks. Consequently, they fail to evaluate model performance for Visually Rich Document Understanding (VRDU), a critical challenge involving complex layouts and dense text. To address this, we introduce DocWeaver, a novel multi-agent pipeline that leverages Large Language Models to automatically generate a new benchmark. The result is MosaicDoc, a large-scale, bilingual (Chinese and English) resource designed to push the boundaries of VRDU. Sourced from newspapers and magazines, MosaicDoc features diverse and complex layouts (including multi-column and non-Manhattan), rich stylistic variety from 196 publishers, and comprehensive multi-task annotations (OCR, VQA, reading order, and localization). With 72K images and over 600K QA pairs, MosaicDoc serves as a definitive benchmark for the field. Our extensive evaluation of state-of-the-art models on this benchmark reveals their current limitations in handling real-world document complexity and charts a clear path for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。