arXiv:2603.29345cs.CL2026-03中稿 · SIGUL 2026

首个开源机器翻译系统评估,为世界语打造高效翻译方案

Open Machine Translation for Esperanto

  • 对比规则系统、编码器-解码器与大模型在六种语言对上的表现
  • NLLB系列在所有方向上表现最优,人类评测中半数偏好其翻译结果
  • 开源代码与最佳模型,推动世界语的开放协作与技术发展

世界语是一种广泛使用的构造语言,以其规则的语法和丰富的构词能力著称。尽管得益于在线社区积累了大量资源,但在现代机器翻译方法中仍相对未被充分探索。本文首次全面评估开源机器翻译系统在世界语中的应用,涵盖规则系统、编码器-解码器模型及不同规模的大语言模型(LLM),覆盖英语、西班牙语、加泰罗尼亚语与世界语之间的六种语言方向。通过多种自动指标与人工评估进行验证。结果显示,NLLB系列在所有语言对中表现最佳,其次为训练的紧凑模型和微调的通用大模型。人工评估确认此趋势,约一半比较中更倾向选择NLLB的翻译结果,但仍有明显错误存在。秉承世界语倡导的开放与国际协作精神,我们公开发布代码及最佳模型。

原文摘要 · Abstract (English)

Esperanto is a widespread constructed language, known for its regular grammar and productive word formation. Besides having substantial resources available thanks to its online community, it remains relatively underexplored in the context of modern machine translation (MT) approaches. In this work, we present the first comprehensive evaluation of open-source MT systems for Esperanto, comparing rule-based systems, encoder-decoder models, and LLMs across model sizes. We evaluate translation quality across six language directions involving English, Spanish, Catalan, and Esperanto using multiple automatic metrics as well as human evaluation. Our results show that the NLLB family achieves the best performance in all language pairs, followed closely by our trained compact models and a fine-tuned general-purpose LLM. Human evaluation confirms this trend, with NLLB translations preferred in approximately half of the comparisons, although noticeable errors remain. In line with Esperanto's tradition of openness and international collaboration, we release our code and best-performing models publicly.

机器翻译世界语开源模型多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。