arXiv:2608.10396cs.CV2026-08

新基准可诊断表格文档结构识别的失败位置,揭示模型深层缺陷。

FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition

论文配图:FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition
图 1 · 摘自论文原文
  • 构建分层诊断基准,从文档到组件逐级评估结构还原能力。
  • 最佳整体准确率83.85%,细粒度结构识别仍低于18%。
  • 适用于评测表格理解模型在复杂布局下的鲁棒性与可解释性。

将表格类文档转化为机器可处理的结构化数据,不仅需要恢复可见内容,还需重建其多层级组织结构。现有基准仅评估整体输出或常规表格网格,聚合得分无法定位结构性错误。我们提出FormStruct-Bench,一个分层且可诊断的基准,支持在文档级及更细粒度组件级评估结构识别,实现性能问题可追溯。为构建可审计的真实标签,我们标注70个可复用模板,并通过保溯源的Director--Artist--Verifier流程扩展为7,000个验证实例;测试集中1,100个与模板无关的实例均经人工复核。评估协议包含五项核心指标和三项结构专项诊断,覆盖页面、模式与组件层级,并按难度、结构约束与视觉退化进行切片分析。在14个支持API调用与本地部署的系统及两个SFT变体上,最高文档级得分为83.85%,但最细粒度结构得分仍低于18%。结果揭示了内容读取与结构层次还原之间显著差距,凸显可靠表格理解所需的能力鸿沟。

原文摘要 · Abstract (English)

Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure that organizes it. However, existing benchmarks evaluate either holistic document outputs or conventional table grids, and their aggregate scores provide little insight into where structural failures occur. We introduce FormStruct-Bench, a hierarchical and diagnostic benchmark that evaluates table-form document structure recognition at both the document level and progressively finer component levels, allowing aggregate performance to be traced back to specific structural failure modes. To construct auditable ground truth at scale, we annotate 70 reusable templates and expand them into 7,000 verified instances through a provenance-preserving Director--Artist--Verifier pipeline; all 1,100 instances in the template-disjoint test set additionally receive human review. Our evaluation protocol uses five primary metrics and three structure-specific diagnostics across page, schema, and component levels, together with slices over difficulty, structural constraints, and visual degradation. Across 14 API-hosted and locally deployable systems plus two SFT variants, the best document-level score reaches 83.85%, whereas the best reported fine-grained structural score remains below 18%. These results reveal a pronounced gap between reading document content and recovering the hierarchy and regional organization required for reliable table-form understanding.

文档理解结构识别评估基准分层诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。