arXiv:2603.04205cs.CV2026-03被引 13

构建真实文档解析的全尺度测试基准,揭示模型在真实场景下的性能短板。

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild

  • 对1355张数字文档进行五类真实场景的物理重建,实现精准对照。
  • 发现模型在扫描、扭曲、光照等真实干扰下性能显著下降。
  • 适合研究鲁棒文档解析与现实世界泛化能力的团队使用。

尽管视觉语言模型(VLMs)在数字文档基准(如OmniDocBench)上表现接近完美,但在不可预测的真实物理世界中的表现仍不明确,原因在于缺乏可控且真实的评估。我们提出Real5-OmniDocBench,首个对OmniDocBench v1.5(共1,355张图像)进行完整、一对一物理重建的基准,覆盖五种关键真实场景:扫描、扭曲、屏幕拍摄、光照变化和倾斜。与以往仅部分采样或缺乏数字对应关系的基准不同,我们的完整真值映射首次实现了性能退化的因素分离,可精准定位失败是源于几何失真、光学伪影还是模型本身局限。该基准为社区设立了新挑战,揭示文档解析中的‘现实差距’远未弥合,并提供诊断工具以指导真正鲁棒的文档智能研发。

原文摘要 · Abstract (English)

While Vision-Language Models (VLMs) achieve near-perfect scores on digital document benchmarks like OmniDocBench, their performance in the unpredictable physical world remains largely unknown due to the lack of controlled yet realistic evaluations. We introduce Real5-OmniDocBench, the first benchmark that performs a full-scale, one-to-one physical reconstruction of the entire OmniDocBench v1.5 (1,355 images) across five critical real-world scenarios: Scanning, Warping, Screen-Photography, Illumination, and Skew. Unlike prior benchmark that either lack digital correspondence or employ partial sampling, our complete ground-truth mapping enables, for the first time, rigorous factor-wise attribution of performance degradation-allowing us to pinpoint whether failures stem from geometric distortions, optical artifacts, or model limitations. Our benchmark establishes a challenging new standard for the community, demonstrating that the 'reality gap' in document parsing is far from closed, and provides a diagnostic tool to guide the development of truly resilient document intelligence.

文档解析真实场景基准测试视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。