针对科学文档跨版本差异检测,提出兼顾布局与结构的智能比对方法。
Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning

- 先对齐后推理:融合空间、内容与结构信息建立元素对应关系。
- 文本/表格/公式/图表差异检测F1达0.903~0.845,显著优于基线。
- 适合期刊编辑、技术文档维护人员快速定位修改内容。
科学文档跨版本差异检测在学术出版与技术文档中至关重要,但因文档包含文本、表格、公式、图表及版式线索等异构元素而面临挑战。现有基于文本序列的方法易丢失版式与结构信息,图像方法则缺乏语义可解释性且对渲染变化敏感。本文提出一种布局感知的异构元素感知框架,将文档分解为语义类型化元素,通过联合建模空间、内容与结构兼容性的对齐优先机制建立跨版本对应关系,并对齐元素对进行类型感知的差异推理。该框架支持统一的变更检测、定位、结构感知分析及匹配评估,覆盖文本、表格、公式和图表。在真实期刊排校流程的科学PDF数据上实验表明,该框架持续优于特定任务基线,文本、表格、公式、图表的检测F1分别为0.903、0.855、0.862、0.845,且在定位精度、结构感知与匹配质量上均有提升。消融与敏感性分析验证了跨版本对齐、类型特异性表示、结构感知推理及兼容性权重设计的有效性。结果表明,该方法为真实编辑生产场景下的科学文档比对提供了鲁棒且可解释的解决方案。
原文摘要 · Abstract (English)
Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challenging because scientific documents are page-structured artifacts containing heterogeneous elements such as text, tables, formulas, figures, and layout cues. Existing text-sequence-based methods often lose layout and structural information, while image-based methods lack semantic interpretability and are sensitive to rendering variation. To address these limitations, this paper proposes a layout-aware heterogeneous element-aware framework for scientific document differencing. The framework decomposes document versions into semantically typed elements, establishes cross-version correspondence through an alignment-first mechanism that jointly models spatial, content, and structural compatibility, and performs type-aware difference reasoning over aligned element pairs. It supports unified change detection, localization, structure-awareness analysis, and alignment/matching evaluation across text, tables, formulas, and figures. Experiments on real-world scientific PDF data from journal production proofreading workflows show that the proposed framework consistently outperforms element-specific baselines. It achieves detection F1 scores of 0.903, 0.855, 0.862, and 0.845 for text, tables, formulas, and figures, respectively, with further improvements in localization, structure awareness, and matching quality. Ablation and sensitivity analyses confirm the effectiveness of cross-version alignment, type-specific representations, structure-aware reasoning, and compatibility-weight design. These results demonstrate that heterogeneous element-aware differencing provides a robust and interpretable solution for scientific document comparison in realistic editorial production scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。