首个面向真实场景的文档理解评测基准,揭示大模型在复杂环境下的脆弱性。
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
- 构建真实拍摄文档数据集,覆盖光照、形变等自然干扰。
- 同一文档四次拍摄对比,暴露模型性能下降超30%。
- 适合关注真实应用中模型鲁棒性的研究者与工程师。
多模态大模型虽显著提升了文档理解能力,但现有评测基准如DocVQA和ChartQA主要基于扫描或数字文档,难以反映真实世界中的复杂挑战,如光照变化与物理形变。本文提出首个针对自然环境下文档理解的评测基准WildDoc,包含大量人工拍摄的真实文档图像,并整合已有基准来源以支持与数字化文档的全面比较。为严格评估模型鲁棒性,每份文档在四种不同条件下各拍摄一次。对前沿多模态大模型在WildDoc上的测试显示,其性能显著下降,暴露出相比传统基准在真实场景中严重缺乏稳健性,凸显了实际文档理解的独特挑战。项目主页:https://bytedance.github.io/WildDoc。
原文摘要 · Abstract (English)
The rapid advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced capabilities in Document Understanding. However, prevailing benchmarks like DocVQA and ChartQA predominantly comprise \textit{scanned or digital} documents, inadequately reflecting the intricate challenges posed by diverse real-world scenarios, such as variable illumination and physical distortions. This paper introduces WildDoc, the inaugural benchmark designed specifically for assessing document understanding in natural environments. WildDoc incorporates a diverse set of manually captured document images reflecting real-world conditions and leverages document sources from established benchmarks to facilitate comprehensive comparisons with digital or scanned documents. Further, to rigorously evaluate model robustness, each document is captured four times under different conditions. Evaluations of state-of-the-art MLLMs on WildDoc expose substantial performance declines and underscore the models' inadequate robustness compared to traditional benchmarks, highlighting the unique challenges posed by real-world document understanding. Our project homepage is available at https://bytedance.github.io/WildDoc.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。