构建首个古希腊语到现代希腊语翻译数据集,提升低资源语言翻译性能。
Ancient Greek to Modern Greek Machine Translation: A Novel Benchmark and Fine-Tuning Experiments on LLMs and NMT Models

- 用网页抓取与多阶段对齐生成13万句古现代希腊语平行语料。
- 微调后模型最高提升10.3点BLEU,Llama-Krikri-8B表现最优。
- 适合希腊语NLP研究者和低资源语言翻译方向从业者。
古希腊语(AG)到现代希腊语(MG)的机器翻译属于低资源任务,受限于缺乏大规模高质量平行数据。本文提出AG-MG平行语料库,包含132,481句对齐句子,源自文学、历史与圣经文本。我们设计了一套新语料构建流程:结合网页抓取的片段级数据,通过多阶段句级对齐与精修过程实现高质量匹配。方法采用VecAlign与LaBSE嵌入,先在人工对齐的子集上微调,再使用Gemini 2.5 Flash进行大模型纠错,确保对齐精度。此外,首次系统评估现代MT模型在该任务的表现,涵盖三种微调策略,对比NMT模型(NLLB、M2M100)与希腊语LLM(Llama-Krikri-8B)。实验显示微调显著优于基线模型,最高提升达+10.3 BLEU点。其中,全参数微调的Llama-Krikri-8B取得最高得分13.16;而QLoRA适配的M2M100-1.2B模型相对增益最大且表现优异。本数据集与模型为希腊语NLP发展提供重要支持。
原文摘要 · Abstract (English)
Machine Translation (MT) for Ancient Greek (AG) to Modern Greek (MG) is a low-resource task, constrained by the lack of large-scale, high-quality parallel data. We address this gap by introducing the AG-MG Parallel Corpus, a new resource containing 132,481 sentence-aligned pairs derived from literary, historical, and biblical texts. We present a novel corpus creation pipeline that combines web-scraped, excerpt-level data with a multi-stage sentence-level alignment, and refinement process. Our method uses VecAlign with LaBSE embeddings, which we first fine-tune on a manually-aligned AG-MG subset, followed by an LLM-based error/misalignment correction phase using Gemini 2.5 Flash to ensure high alignment quality. Furthermore, we provide the first comprehensive benchmark of modern MT models on this task, evaluating three fine-tuning strategies across NMT models (NLLB, M2M100) and a Greek LLM (Llama-Krikri-8B). Our experiments show that fine-tuning yields significant improvements over base models, increasing performance by up to +10.3 BLEU points. Specifically, full-parameter fine-tuning of Llama-Krikri-8B achieves the highest overall performance with a BLEU score of 13.16, while the QLoRA-adapted M2M100-1.2B model demonstrates the largest relative gains and highly competitive results. Our dataset and models represent a significant contribution to Greek NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。