评测真实拍摄文档的解析与翻译,发现模型性能下降超18%。
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
- 构建真实拍摄文档的多领域评测集,覆盖8种翻译场景。
- 真实拍摄使主流模型解析准确率下降18%,专用模型下降25%。
- 揭示现有模型在复杂拍摄条件下的脆弱性,适合评估鲁棒性。
多模态大模型(MLLMs)推动了端到端文档解析与翻译的发展。然而,当前基准如OmniDocBench和DITrans主要基于扫描或数字生成文档,难以反映真实拍摄中的几何畸变和光照变化等挑战。为此,我们提出DocPTBench,一个专为拍摄文档解析与翻译设计的综合性评测基准。该数据集包含超过1,300张高分辨率的真实拍摄文档,覆盖多个领域,提供8种翻译场景,并配有精心人工验证的解析与翻译标注。实验表明,从数字文档转向真实拍摄文档,主流MLLM在端到端解析上平均准确率下降18%,翻译下降12%;专用文档解析模型平均下降25%。这一显著性能差距凸显了真实拍摄条件带来的独特挑战,揭示了现有模型的有限鲁棒性。数据集与代码已开源:https://github.com/Topdu/DocPTBench。
原文摘要 · Abstract (English)
The advent of Multimodal Large Language Models (MLLMs) has unlocked the potential for end-to-end document parsing and translation. However, prevailing benchmarks such as OmniDocBench and DITrans are dominated by pristine scanned or digital-born documents, and thus fail to adequately represent the intricate challenges of real-world capture conditions, such as geometric distortions and photometric variations. To fill this gap, we introduce DocPTBench, a comprehensive benchmark specifically designed for Photographed Document Parsing and Translation. DocPTBench comprises over 1,300 high-resolution photographed documents from multiple domains, includes eight translation scenarios, and provides meticulously human-verified annotations for both parsing and translation. Our experiments demonstrate that transitioning from digital-born to photographed documents results in a substantial performance decline: popular MLLMs exhibit an average accuracy drop of 18% in end-to-end parsing and 12% in translation, while specialized document parsing models show significant average decrease of 25%. This substantial performance gap underscores the unique challenges posed by documents captured in real-world conditions and reveals the limited robustness of existing models. Dataset and code are available at https://github.com/Topdu/DocPTBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。