首个开源韩文古籍汉文处理平台,助力古籍翻译与理解
HERITAGE: An End-to-End Web Platform for Processing Korean Historical Documents in Hanja
- 基于汉文语言模型实现标点恢复、命名实体识别与机器翻译
- 提供韩/英双语翻译及字符级释义交互词典,提升可读性
- 支持专家修正模型输出,加速古籍数字化进程
韩国历史文献是宝贵的文化遗产,但理解这些文献需深厚的汉文知识。汉文为20世纪前韩国使用的古文字,源自古代汉语但在韩国历经数百年演变。现代韩、中两国读者均难以直接理解,尽管已有部分韩文和英文译本,仍需专业知识,导致多数文献未被翻译。为此,我们提出HERITAGE——首个开源的汉文NLP工具包,用于处理韩文古籍中的汉文文本。HERITAGE是一个基于网页的平台,通过汉文语言模型提供三项关键任务的模型预测:标点恢复、命名实体识别和机器翻译(MT)。平台还配备交互式词典,提供汉文字符的现代韩语读音及字符级英文释义。该平台有两个用途:一是让普通用户通过模型输出和词典获得文献大致理解,特别是韩文和英文的翻译结果;二是由于模型输出不完美,汉文专家可对其进行修订,生成更准确的标注与翻译。此举将显著提升翻译效率,有望使大部分历史文献被翻译为现代语言,降低对未开发韩文古籍的理解门槛。
原文摘要 · Abstract (English)
While Korean historical documents are invaluable cultural heritage, understanding those documents requires in-depth Hanja expertise. Hanja is an ancient language used in Korea before the 20th century, whose characters were borrowed from old Chinese but had evolved in Korea for centuries. Modern Koreans and Chinese cannot understand Korean historical documents without substantial additional help, and while previous efforts have produced some Korean and English translations, this requires in-depth expertise, and so most of the documents are not translated into any modern language. To address this gap, we present HERITAGE, the first open-source Hanja NLP toolkit to assist in understanding and translating the unexplored Korean historical documents written in Hanja. HERITAGE is a web-based platform providing model predictions of three critical tasks in historical document understanding via Hanja language models: punctuation restoration, named entity recognition, and machine translation (MT). HERITAGE also provides an interactive glossary, which provides the character-level reading of the Hanja characters in modern Korean, as well as character-level English definition. HERITAGE serves two purposes. First, anyone interested in these documents can get a general understanding from the model predictions and the interactive glossary, especially MT outputs in Korean and English. Second, since the model outputs are not perfect, Hanja experts can revise them to produce better annotations and translations. This would boost the translation efficiency and potentially lead to most of the historical documents being translated into modern languages, lowering the barrier on unexplored Korean historical documents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。