arXiv:2509.00909cs.IR2025-09中稿 · as a demo paper at…

解决法律类书籍的深度章节结构解析难题,提升段落边界识别精度。

HiPS: Hierarchical PDF Segmentation of Doctrinal Legal Books

  • 结合目录元数据与大模型优化的双通道解析流程
  • 在9812个手工标注标题上实现高精度层次重建
  • 适合法律文献数字化与司法文本分析研究者

PDF解析器在页面级版式理解方面已有进步,但对于深层结构化书籍,仍难以可靠恢复文档级的章节层级关系:许多系统仅能识别页内标题角色,假设层级较浅,或依赖高质量的PDF标签与目录元数据,且公开的深层书籍层级金标准数据稀缺。本文提出HiPS,用于教义性法律书籍的分层PDF分割,并做出两项主要贡献:首先,发布了一个包含49本开放获取法律书籍的金标准基准数据集,涵盖9,812个手工标注的标题、层级关系和页码锚点,支持标题检测、层级重构与章节边界判定的评估;其次,提出互补的分割流水线:一种基于目录元数据的解析器,适用于元数据完整的书籍;另一种无目录的LLM精炼流水线,融合OCR空白区域线索、XML排版信息与局部上下文。在广泛对比开源解析器及多模态/大模型基线时,有目录的流水线在元数据完整时表现最佳,而无目录的LLM精炼流水线在元数据缺失或噪声环境下显著提升标题精确率、深层结构恢复能力与边界质量。

原文摘要 · Abstract (English)

PDF parsers have recently improved on page-level layout understanding. However, recovering a document-global section hierarchy with reliable boundaries remains brittle for deeply structured books: many systems expose only page-local heading roles, assume shallow depth, or rely on high-quality PDF tags or Table of Contents (TOC) metadata, and public gold-standard data for deep book hierarchies is scarce. We present HiPS for hierarchical PDF segmentation of doctrinal legal books and make two main contributions. First, we release a gold-standard benchmark of 49 open-access law books with 9,812 manually curated headings, hierarchy levels, and page anchors, enabling evaluation of title detection, hierarchy reconstruction, and section boundary assignment. Second, we introduce complementary segmentation pipelines: a TOC-based parser for books with reliable outline metadata and a TOC-free LLM-refined pipeline that combines OCR whitespace cues, XML typography, and local context. Across a broad comparison against open-source parsers and multimodal/LLM baselines, the TOC-based pipeline is strongest when metadata is complete, while the LLM-refined pipeline improves heading precision, deep-level recovery, and boundary quality when metadata is missing or noisy.

法律文本文档解析大模型层级结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。