arXiv:2511.05239cs.CL2025-11Conference of the …

用标注系统模拟古籍中日文翻译,提升低资源下翻译质量。

Translation via Annotation: A Computational Study of Translating Classical Chinese into Japanese

  • 将古籍注释视为序列标注任务,融入现代语言技术。
  • 引入辅助中文NLP任务,在低资源下提升标注精度。
  • 适合研究古典文献数字化与跨语言注释的学者。

古代人通过在汉字周围添加注释的方式将古典汉语翻译为日语。本文将这一过程抽象为序列标注任务,并融入现代语言技术。该研究面临资源匮乏问题,我们通过基于大模型的注释流程构建新数据集以缓解此问题。实验表明,在低资源条件下,引入辅助中文NLP任务能有效提升序列标注任务的训练效果。同时评估了大语言模型在此任务上的表现:尽管在直接机器翻译中得分较高,但本文方法可作为补充,显著提升字符注释质量。

原文摘要 · Abstract (English)

Ancient people translated classical Chinese into Japanese using a system of annotations placed around characters. We abstract this process as sequence tagging tasks and fit them into modern language technologies. The research on this annotation and translation system faces a low resource problem. We alleviate this problem by introducing an LLM-based annotation pipeline and constructing a new dataset from digitized open-source translation data. We show that in the low-resource setting, introducing auxiliary Chinese NLP tasks enhances the training of sequence tagging tasks. We also evaluate the performance of Large Language Models (LLMs) on this task. While they achieve high scores on direct machine translation, our method could serve as a supplement to LLMs to improve the quality of character's annotation.

古典翻译序列标注低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。