为濒危语言阿罗马尼亚语构建了首个大规模翻译系统。
Dialectal and Low-Resource Machine Translation for Aromanian
- 构建7.9万句对的阿罗马尼亚语-罗马尼亚语平行语料库
- 开发并对比多个适配阿罗马尼亚语的神经机器翻译模型
- 适合语言保护与低资源语种研究者使用
本文介绍了一个支持英语、罗马尼亚语和阿罗马尼亚语(一种濒危东罗曼语)的神经机器翻译系统的构建过程。主要贡献有二:(1) 构建了迄今为止规模最大的阿罗马尼亚语-罗马尼亚语平行语料库,包含79,000句对;(2) 开发并比较了多个针对阿罗马尼亚语优化的机器翻译模型。为实现这一目标,我们引入了一套辅助工具,包括一个语言无关的句子嵌入模型用于文本挖掘和自动评估,以及一个支持不同书写标准的变音符号转换系统。该研究在计算语言学和语言保护方面均做出贡献,为历史上的低资源语言建立了关键资源。所有数据集、训练模型及相关工具均已公开:https://huggingface.co/aronlp 和 https://arotranslate.com。
原文摘要 · Abstract (English)
This paper presents the process of building a neural machine translation system with support for English, Romanian, and Aromanian - an endangered Eastern Romance language. The primary contribution of this research is twofold: (1) the creation of the most extensive Aromanian-Romanian parallel corpus to date, consisting of 79,000 sentence pairs, and (2) the development and comparative analysis of several machine translation models optimized for Aromanian. To accomplish this, we introduce a suite of auxiliary tools, including a language-agnostic sentence embedding model for text mining and automated evaluation, complemented by a diacritics conversion system for different writing standards. This research brings contributions to both computational linguistics and language preservation efforts by establishing essential resources for a historically under-resourced language. All datasets, trained models, and associated tools are public: https://huggingface.co/aronlp and https://arotranslate.com
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。