针对历史报纸版面复杂难解析问题,提出多模态语义组装新方法。
PereStruct: Multimodal Semantic Assembly for Robust Historical Document Parsing

- 结合视觉检测与语义相似度,重建文章结构
- 在块到文章映射上达到0.904的F1分数
- 适合历史文献数字化与古籍智能处理研究者
解析具有复杂非标准排版的历史文献仍是大规模档案数字化的核心瓶颈。与现代排版不同,历史报纸存在严重物理退化和高度不规则的页面结构,使最先进的视觉语言模型也难以应对,带来严重的分布外挑战。我们提出一个专为历史报纸设计的自动化解析流水线,其文档以复杂的多栏布局为特征。该方法结合在1,426张人工标注扫描页上微调的YOLO架构进行版面分析与区块检测,以及一种新颖的语义组装模块,通过联合建模词法-语义相似性(TF-IDF)、微调后的YOLO视觉嵌入和几何布局约束来重构文章。这种多模态融合实现最优性能,在块到文章映射上获得0.904的F1分数。端到端评估显示,相较于视觉语言模型(Qwen3.6-35B-A3B 和 Qwen3.6-Plus),PereStruct的生成质量显著更高(BLEU约0.96对比0.34),验证了模块化架构在复杂历史版面中优于通用VLM。为支持可复现性并推动该领域研究,我们发布包含599个标注页的训练语料库及93个经专家验证真实标签的PereStruct基准数据集。该框架为复杂档案材料的高保真数字化与语义重构奠定了坚实基础。
原文摘要 · Abstract (English)
Parsing historical documents with complex, non-standard layouts remains a fundamental bottleneck in large-scale archival digitization. Unlike modern typography, historical newspapers exhibit severe physical degradation and highly irregular page structures that confound even state-of-the-art vision-language models, presenting severe out-of-distribution challenges. We address this gap with an automated pipeline specifically designed for parsing historical newspapers, documents characterized by particularly intricate multi-column layouts. Our approach combines a fine-tuned YOLO architecture for layout analysis and block detection, trained on 1,426 fully human-annotated scanned pages, with a novel semantic assembly module that reconstructs articles by jointly modeling lexical-semantic similarity via TF-IDF, visual embeddings from our fine-tuned YOLO, and geometric layout constraints. This multi-modal integration yields state-of-the-art performance, achieving an F1 score of 0.904 on block-to-article mapping. Notably, end-to-end evaluation against vision-language models (Qwen3.6-35B-A3B and Qwen3.6-Plus) demonstrates that PereStruct achieves substantially higher fidelity (BLEU approximately 0.96 vs 0.34), validating that modular architectures excel where generic VLMs fail on complex historical layouts. To support reproducibility and advance research in this domain, we release both the training corpus of 599 annotated pages and a curated PereStruct benchmark of 93 pages with expert-verified ground-truth block-to-article mappings. This framework establishes a robust foundation for high-fidelity digitization and semantic reconstruction of complex archival materials.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。