arXiv:2607.21936cs.CL2026-07ACL

用外部知识增强大模型,提升古籍文本中人名地名的还原准确率。

Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models

论文配图:Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models
图 1 · 摘自论文原文
  • 结合预训练模型与检索到的外部历史信息进行文本修复
  • 在韩文古籍上恢复人名地名效果显著优于基线方法
  • 适合历史学者快速解析破损文献,提升研究效率

历史文献是宝贵的知识遗产,但常因物理损毁导致字迹模糊。现有基于掩码语言建模的修复方法虽能利用局部上下文,却难以还原依赖外部历史知识的专有名词。为此,我们提出一种基于检索增强生成(RAG)的大语言模型框架ARI,融合预训练模型的隐含知识与显式检索的外部上下文,有效解决上下文相关专有名词推断难题。在韩文历史文献上的大量实验表明,该方法在恢复通用字符和专有名词方面均显著优于基线模型。此外,包括专家评估在内的全面评测证实,ARI可作为领域专家的实用工具,有望加速历史文献分析进程。

原文摘要 · Abstract (English)

Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing restoration methods based on masked language modeling effectively utilize local context, they struggle to restore named entities that require external historical knowledge. To address this limitation, we introduce a novel framework for historical document restoration that leverages large language models with retrieval-augmented generation (RAG). By combining the implicit knowledge of pre-trained LLMs with explicitly retrieved external context, our model ARI effectively mitigates the challenge of inferring context-dependent proper nouns. Extensive experiments on Korean historical documents demonstrate that our approach significantly outperforms baselines, achieving substantial gains in restoring both general characters and named entities. Furthermore, comprehensive evaluations including expert assessments confirm that ARI serves as a practical tool for domain experts, promising to accelerate the analysis of historical records.

古籍修复大模型知识增强RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。