构建首个法语版OLDI种子语料库,助力法国地区语言平行语料收集。
A French Version of the OLDI Seed Corpus
- 融合机器翻译与母语者校对,构建法语语料库。
- 处理技术术语与维基用户内容混杂的复杂文本。
- 为资源匮乏的法国地方语言提供关键数据支持。
我们提交了首个法语版OLDI种子语料库,作为WMT 2025开放语言数据倡议(OLDI)共享任务的一部分。该语料库通过多套机器翻译系统生成,并借助自建界面由合格母语者进行后编辑完成。源数据结合了高度专业的百科术语与维基用户生成内容特有的风格不规则性,带来独特翻译挑战。该法语语料库并非终点,而是作为关键枢纽资源,旨在促进法国地区语言平行语料的收集。
原文摘要 · Abstract (English)
We present the first French partition of the OLDI Seed Corpus, our submission to the WMT 2025 Open Language Data Initiative (OLDI) shared task. We detail its creation process, which involved using multiple machine translation systems and a custom-built interface for post-editing by qualified native speakers. We also highlight the unique translation challenges presented by the source data, which combines highly technical, encyclopedic terminology with the stylistic irregularities characteristic of user-generated content taken from Wikipedia. This French corpus is not an end in itself, but is intended as a crucial pivot resource to facilitate the collection of parallel corpora for the under-resourced regional languages of France.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。