arXiv:2608.28170cs.CLcs.AI2026-08

用语言模型修复古籍缺失文字,辅助学者研究。

Text Restoration of Ancient Documents with Language Models

论文配图:Text Restoration of Ancient Documents with Language Models
图 1 · 摘自论文原文
  • 根据文档结构和缺字长度选择不同语言模型,适配真实修复场景。
  • 模型表现受文本类型和缺字长度影响显著,无法完全自动化。
  • 首次系统分析公式化与非公式化内容的修复效果差异。

本研究探讨利用语言模型恢复因物理破损导致的古籍文字缺失的可行性。通过构建多种模拟真实情况的修复场景,针对不同场景选用适配的语言模型架构,并提出多种解码策略以缓解缺字边界与模型分词方案之间的不一致问题。结果表明,文本修复难以完全自动化,但可作为辅助工具帮助古文献学家开展工作;模型性能随需修复部分的结构特征及缺失字符长度而显著变化。本研究首次系统分析了公式化与非公式化内容的修复表现差异,以及缺字长度感知对修复效果的影响,这两者均为古文献学家手动修复中的常见挑战。通过在不同设置下对多种模型进行定性与定量比较,本研究为开发辅助修复工具提供了实用指南。

原文摘要 · Abstract (English)

Purpose - This study investigates the feasibility of restoring missing text caused by physical lacunae in damaged ancient manuscripts using language models. Methodology - The study proposes different scenarios to replicate real-world conditions. Language models of different architectures are applied according to their suitability to each scenario. We also propose several decoding strategies that further enhance performance and address the discrepancy between lacuna boundaries and the models' tokenization schemes. Findings - The results reveal that text restoration of these documents cannot be fully automated, but it can serve as a useful tool to assist paleographers in their work. Model performance varies greatly depending on which structural part of the document needs to be restored and whether the character length of missing text is available. Originality - This is the first study and to analyze model performance on formulaic and non-formulaic content and the impact of lacuna length awareness in manuscript restoration. Both are recurring challenges in paleographers' manual restoration work. Through systematic comparison and both qualitative and quantitative analysis of different models' performance under varying settings, this study offers a guideline for developing assistive tools to support paleographers.

古籍修复语言模型文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。