arXiv:2603.17265cs.CVcs.CL2026-03被引 1

提出文档布局错误检测基准,精准识别模型结构推理缺陷

LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis

  • 定义8类标准错误类型,构建可量化模拟的错误注入机制
  • 实测多模态模型在结构推理上普遍存在明显弱点
  • 适合评估文档理解模型的逻辑一致性与鲁棒性

大语言模型和多模态模型虽提升文档版面分析能力,但区域合并、拆分、遗漏等结构性错误仍普遍存在。传统重叠率指标(如IoU、mAP)无法捕捉此类逻辑不一致。为此,我们提出布局错误检测(LED)基准,超越表面准确率,评估版面分析预测中的结构推理能力。LED定义八类标准化错误类型(缺失、幻觉、尺寸错误、拆分、合并、重叠、重复、误分类),并提供量化规则与真实错误模拟的注入算法。基于此构建LED-Dataset,设计三项评估任务:文档级错误检测、文档级错误类型分类、元素级错误类型分类。对前沿多模态模型的实验表明,LED能实现细粒度、可解释的结构理解评估,揭示不同模态与架构下的明显短板。总体而言,LED建立了统一且可解释的基准,用于诊断文档理解模型的结构鲁棒性与推理能力。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) and Large Multimodal Models (LMMs) have improved Document Layout Analysis (DLA), yet structural errors such as region merging, splitting, and omission remain persistent. Conventional overlap-based metrics (e.g., IoU, mAP) fail to capture such logical inconsistencies. To overcome this limitation, we propose Layout Error Detection (LED), a benchmark that evaluates structural reasoning in DLA predictions beyond surface-level accuracy. LED defines eight standardized error types (Missing, Hallucination, Size Error, Split, Merge, Overlap, Duplicate, and Misclassification) and provides quantitative rules and injection algorithms for realistic error simulation. Using these definitions, we construct LED-Dataset and design three evaluation tasks: document-level error detection, document-level error-type classification, and element-level error-type classification. Experiments with state-of-the-art multimodal models show that LED enables fine-grained and interpretable assessment of structural understanding, revealing clear weaknesses across modalities and architectures. Overall, LED establishes a unified and explainable benchmark for diagnosing the structural robustness and reasoning capability of document understanding models.

文档分析错误检测多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。