评测15个视觉语言模型在法语文档转Markdown上的表现,聚焦手写和复杂版式。
Benchmarking Vision-Language Models for French PDF-to-Markdown Conversion
- 基于模型分歧采样构建法语文档基准,覆盖手写表单、复杂布局等挑战场景。
- 最强商业模型在手写和表单上表现更稳,开源模型在标准印刷体中仍具竞争力。
- 设计针对性测试用例,忽略仅影响排版的无关差异,更贴近下游应用需求。
本报告评估了近期视觉语言模型(VLMs)在处理具有挑战性的法语文档时的PDF转Markdown性能。文档解析是检索增强生成(RAG)流程中的关键步骤,文本与版式错误会传递至下游检索与定位任务。现有基准多侧重英语或中文,容易过度惩罚对下游任务无实质影响的格式与线性化选择(如换行、列表分割、表格不同呈现方式)。我们构建了一个以法语为核心的基准,从60,000份文档中通过模型分歧采样选取困难页面,涵盖手写表单、复杂布局、密集表格及图文混排页面。评估采用单元测试风格的检查项,针对具体失效模式(文本存在性、阅读顺序、局部表格约束),并结合类别特异性归一化策略,剔除仅由表现形式引起的差异。在15个模型中,最强的专有模型在手写与表单任务上表现出显著更高的鲁棒性,而若干开源权重系统在标准印刷布局上仍保持竞争力。
原文摘要 · Abstract (English)
This report evaluates PDF-to-Markdown conversion using recent Vision-Language Models (VLMs) on challenging French documents. Document parsing is a critical step for Retrieval-Augmented Generation (RAG) pipelines, where transcription and layout errors propagate to downstream retrieval and grounding. Existing benchmarks often emphasize English or Chinese and can over-penalize benign formatting and linearization choices (e.g., line breaks, list segmentation, alternative table renderings) that are largely irrelevant for downstream use. We introduce a French-focused benchmark of difficult pages selected via model-disagreement sampling from a corpus of 60{,}000 documents, covering handwritten forms, complex layouts, dense tables, and graphics-rich pages. Evaluation is performed with unit-test-style checks that target concrete failure modes (text presence, reading order, and local table constraints) combined with category-specific normalization designed to discount presentation-only variance. Across 15 models, we observe substantially higher robustness for the strongest proprietary models on handwriting and forms, while several open-weights systems remain competitive on standard printed layouts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。