提出新基准LED,专治文档布局分析中的结构错误
LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
- 定义8类标准结构错误,设计三类诊断任务
- 在多个大模型上验证,发现传统指标无法察觉的缺陷
- 适合关注布局理解鲁棒性的研究者和开发者
近年来,基于大语言模型和多模态模型的文档布局分析取得了显著进展,但在处理区域合并、分割和内容缺失等关键结构性错误方面仍存在挑战。传统评估指标如IoU和mAP主要关注空间重叠,难以捕捉此类错误。为此,我们提出布局错误检测(LED)基准,定义了八种标准化错误类型,并设计三类互补任务:错误存在性检测、错误类型分类与逐元素错误分类。同时构建了LED-Dataset,一个基于真实布局模型经验分布生成的合成数据集,通过注入逼真的结构性错误。在多种大模型上的实验表明,LED能有效区分模型在结构理解上的能力差异,揭示出传统指标未暴露的模态偏差与性能权衡。
原文摘要 · Abstract (English)
Recent advancements in Document Layout Analysis through Large Language Models and Multimodal Models have significantly improved layout detection. However, despite these improvements, challenges remain in addressing critical structural errors, such as region merging, splitting, and missing content. Conventional evaluation metrics like IoU and mAP, which focus primarily on spatial overlap, are insufficient for detecting these errors. To address this limitation, we propose Layout Error Detection (LED), a novel benchmark designed to evaluate the structural robustness of document layout predictions. LED defines eight standardized error types, and formulates three complementary tasks: error existence detection, error type classification, and element-wise error type classification. Furthermore, we construct LED-Dataset, a synthetic dataset generated by injecting realistic structural errors based on empirical distributions from DLA models. Experimental results across a range of LMMs reveal that LED effectively differentiates structural understanding capabilities, exposing modality biases and performance trade-offs not visible through traditional metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。