揭示文档解析器的结构脆弱性,提出可定位故障根源的审计框架。
How Do Document Parsers Break? Auditing Structural Vulnerability in Document Intelligence

- 设计输出级审计框架ProSA,分离探测、靶向与诊断流程。
- 结构损失率指标与OCR不稳定性相关性达0.916,显著优于面积偏差。
- 能区分遮挡与拓扑破坏路径,适合评估文档智能系统可靠性。
文档版面分析(DLA)流水线为信息检索增强生成、长文档问答等系统提供结构化页面表示,但其鲁棒性评估仍以区域为中心。我们识别出这一“足迹偏差”,提出轻量级输出级审计框架ProSA,实现受控探测、策略驱动靶向与结构感知诊断的解耦。ProSA结合块级结构损失率(B-SLR)、粒度感知暴露描述符和路径归因,分析结构身份丢失位置、失效暴露粒度及故障传播路径。在MinerU和PP-StructureV3上对1000页进行测试,受影响区域与扰动引起的OCR不稳定性相关性较弱(R²=0.384/0.110),而B-SLR与其高度一致(R²=0.727/0.916)。暴露描述符进一步区分了遮挡主导与拓扑主导路径,且匹配足迹的结构探针导致的下游问答/检索退化远超面积匹配的擦除。结果推动DLA鲁棒性评估从基于足迹的应力测试转向结构感知的脆弱性审计。
原文摘要 · Abstract (English)
Document Layout Analysis (DLA) pipelines provide structured page representations for retrieval-augmented generation, long-document question answering, and other document intelligence systems, yet their robustness evaluation remains largely area-centric. We identify this Footprint Bias and propose ProSA, a lightweight output-level auditing framework that decouples controlled probing, policy-driven targeting, and structure-aware diagnosis. ProSA combines Block-level Structural Loss Rate (B-SLR), granularity-aware exposure descriptors, and pathway attribution to analyze where structural identity is lost, at what exposure granularity failures emerge, and how failures propagate. Across MinerU and PP-StructureV3 on 1,000 pages, affected area weakly tracks perturbation-induced OCR instability (R^2=0.384/0.110), whereas B-SLR aligns much more closely with it (R^2=0.727/0.916). Exposure descriptors further separate occlusion- and topology-dominant pathways, while matched-footprint structural probes cause much larger downstream QA/retrieval degradation compared to area-matched erasure. These results shift DLA robustness evaluation from footprint-based stress testing toward structure-aware vulnerability auditing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。