arXiv:2507.20111cs.CLcs.AI2025-07被引 1

用AI生成高质量古英语文本,助力濒危语言复兴

AI-Driven Generation of Old English: A Framework for Low-Resource Languages

  • 采用双代理流程分离内容生成与翻译任务
  • 英文到古英语翻译BLEU得分从26提升至65以上
  • 适合语言保护研究者与AI跨领域应用者

保护古代语言对理解人类文化和语言遗产至关重要,但古英语仍严重缺乏资源,限制了其在现代自然语言处理中的应用。我们提出一种可扩展框架,利用先进大语言模型生成高质量古英语文本。方法结合低秩适应(LoRA)参数高效微调、通过反向翻译实现的数据增强,以及分离内容生成(英文)与翻译(古英语)的双代理流水线。自动评估指标(BLEU、METEOR、CHRF)显示显著提升,英文到古英语翻译的BLEU得分从26提高至65以上。专家人工评估也证实生成文本具有高语法准确性和风格一致性。该方法不仅扩展了古英语语料库,还为其他濒危语言的复兴提供了可复用的技术蓝图,有效融合人工智能创新与文化保护目标。

原文摘要 · Abstract (English)

Preserving ancient languages is essential for understanding humanity's cultural and linguistic heritage, yet Old English remains critically under-resourced, limiting its accessibility to modern natural language processing (NLP) techniques. We present a scalable framework that uses advanced large language models (LLMs) to generate high-quality Old English texts, addressing this gap. Our approach combines parameter-efficient fine-tuning (Low-Rank Adaptation, LoRA), data augmentation via backtranslation, and a dual-agent pipeline that separates the tasks of content generation (in English) and translation (into Old English). Evaluation with automated metrics (BLEU, METEOR, and CHRF) shows significant improvements over baseline models, with BLEU scores increasing from 26 to over 65 for English-to-Old English translation. Expert human assessment also confirms high grammatical accuracy and stylistic fidelity in the generated texts. Beyond expanding the Old English corpus, our method offers a practical blueprint for revitalizing other endangered languages, effectively uniting AI innovation with the goals of cultural preservation.

古英语AI生成语言复兴LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。