构建乌克兰语-英语民间故事双语语料库,助力低资源语言翻译
Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages
- 基于现有英译本与新增翻译,构建乌克兰语-英语平行语料库
- 语料库实现词与句级对齐,适合机器翻译模型训练
- 专为民间故事领域设计,兼顾人文与机器翻译需求
民间故事在语言和文化上极具丰富性,是理解源语言的重要窗口。历史上,民间故事翻译仅依赖人工,导致译本数量稀少,限制了对文化传统与习俗的认知。本文基于已有英文译本并补充新译文,创建了一个新的乌克兰语-英语民间故事平行语料库。提出一种结合领域特性的构建与增强方法,考虑该领域的特殊性以及人类与机器翻译目的的差异。语料库实现词与句级对齐,确保语义精准,专为机器翻译模型训练优化。
原文摘要 · Abstract (English)
Folktales are linguistically very rich and culturally significant in understanding the source language. Historically, only human translation has been used for translating folklore. Therefore, the number of translated texts is very sparse, which limits access to knowledge about cultural traditions and customs. We have created a new Ukrainian-To-English parallel corpus of familiar Ukrainian folktales based on available English translations and suggested several new ones. We offer a combined domain-specific approach to building and augmenting this corpus, considering the nature of the domain and differences in the purpose of human versus machine translation. Our corpus is word and sentence-aligned, allowing for the best curation of meaning, specifically tailored for use as training data for machine translation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。