arXiv:2607.10400cs.CVcs.AI2026-07中稿 · COLM

构建可控合成数据集,揭示视觉文档模型在长文本下的三大隐藏缺陷

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

论文配图:SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
图 1 · 摘自论文原文
  • 通过组合设计独立控制文档长度、版式、模态等变量
  • 发现模型随文档变长性能骤降,中间区域最难理解,图表解析能力崩溃
  • 适合研究长文本视觉理解鲁棒性的学者和模型开发者

视觉语言模型(VLMs)在DocVQA、ChartQA和MMLongBench-Doc等文档理解基准上表现强劲。然而,真实文档同时包含长度、版式复杂度、模态混合和问题难度等多种因素,导致模型失败原因难以归因。本文提出SynthDocBench,一个完全合成的长上下文视觉文档理解基准,系统性地控制文档长度、布局结构、模态组成和问题类型。该基准采用组合设计,各因素独立变化,实现对模型行为的受控分析。文档通过大语言模型流水线生成,涵盖六种版式原型,40%随机覆盖以防止模型利用虚假相关性。相比现有基准,SynthDocBench具有更长的文本长度和更丰富的结构多样性。评估七个前沿VLM后,我们发现三种现有基准无法揭示的失效模式:模型性能随文档长度急剧下降;五种模型对文档中段最敏感(中间三分之一最困难);五种模型呈现负向‘早-晚’趋势(最陡下降达8.3个百分点);以及长文档中图表理解能力崩溃。结果表明,当前模型可能过度拟合基准噪声,而非真正掌握鲁棒的长上下文视觉文档理解能力。

原文摘要 · Abstract (English)

Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question difficulty, which makes it difficult to attribute model failures to specific causes. We introduce SynthDocBench, a fully synthetic benchmark for long-context visual document understanding that systematically controls factors including document length, layout structure, modality composition, and question type. The benchmark is constructed using a combinatorial design, each factor is varied independently across generated documents, enabling controlled analysis of model behavior. Documents are generated end to end using an LLM pipeline across six layout archetypes, with a 40 percent random override to prevent models from exploiting spurious correlations. Additionally, SynthDocBench spans long-context documents with substantially greater length and structural diversity than existing benchmarks. Evaluating seven frontier VLMs, we uncover three failure modes that existing benchmarks cannot surface: sharp degradation with document length, a systematic positional sensitivity in which the middle third of a document is hardest for five of six models and five of six models show a negative Early-to-Late trend (steepest decline: 8.3 percentage points), and breakdown of chart comprehension in long-document settings. These results suggest that current models may be overfitting to benchmark artifacts rather than achieving robust long-context visual document understanding.

视觉理解长文本合成数据模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。