arXiv:2505.18683cs.CL2025-05ACL被引 3

Tulun让小语种翻译更准,无需调参就能用术语库提升专业领域翻译质量。

TULUN: Transparent and Adaptable Low-resource Machine Translation

  • 结合神经翻译与大模型后编辑,用术语库引导翻译优化。
  • 在医疗和救灾场景中,对Tetun和Bislama提升16.90~22.41点ChrF++。
  • 适合非技术用户和小团队使用,支持协作式术语管理。

支持低资源语言的机器翻译系统在专业领域表现不佳。尽管已有多种领域自适应方法,但通常需模型微调,对非技术人员和小型组织不友好。为此,我们提出Tulun——一种术语感知的通用翻译方案,融合神经机器翻译与基于大语言模型的后编辑,由现有术语表和翻译记忆库引导。开源网页平台使用户可轻松创建、编辑并复用术语资源,推动人机协作翻译流程,尊重领域知识的同时提升翻译准确率。评估显示,该系统在真实世界与基准测试中均有效:在Tetun和Bislama的医疗与救灾任务中,相比基线系统提升16.90~22.41 ChrF++;在FLORES数据集上,覆盖六种低资源语言,平均优于NLLB-54B 2.8 ChrF点,超越独立的MT与LLM方法。

原文摘要 · Abstract (English)

Machine translation (MT) systems that support low-resource languages often struggle on specialized domains. While researchers have proposed various techniques for domain adaptation, these approaches typically require model fine-tuning, making them impractical for non-technical users and small organizations. To address this gap, we propose Tulun, a versatile solution for terminology-aware translation, combining neural MT with large language model (LLM)-based post-editing guided by existing glossaries and translation memories. Our open-source web-based platform enables users to easily create, edit, and leverage terminology resources, fostering a collaborative human-machine translation process that respects and incorporates domain expertise while increasing MT accuracy. Evaluations show effectiveness in both real-world and benchmark scenarios: on medical and disaster relief translation tasks for Tetun and Bislama, our system achieves improvements of 16.90-22.41 ChrF++ points over baseline MT systems. Across six low-resource languages on the FLORES dataset, Tulun outperforms both standalone MT and LLM approaches, achieving an average improvement of 2.8 ChrF points over NLLB-54B.

机器翻译低资源语言术语管理大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。