构建多语言跨领域翻译质量评估数据集,支持众包扩展。
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
- 8种非英语语言手工构建,覆盖多语种、多领域、多段落文本
- 相比现有数据集,领域更广且适合非专家参与翻译任务
- 推动跨语言众包协作,支持任意书面语言的平行语料收集
BOUQuET 是一个多语言、多中心、多领域/语域的数据集、基准测试与开放协作倡议。该数据集以8种非英语语言手工构建,这些语言均为使用最广泛的语言,具有作为枢纽语言的潜力,可提升翻译准确性。数据集设计注重多语言特征的代表性,并突破句子级限制,采用不同长度的段落组织形式。相较于已有机器翻译数据集,BOUQuET在领域覆盖上更为广泛,同时简化了非专业人员的翻译任务。因此,本研究特别适用于众包扩展,现已发起号召,旨在收集涵盖任何书面语言的多向平行语料库。
原文摘要 · Abstract (English)
BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages. Each of these source languages are representative of the most widely spoken ones and therefore they have the potential to serve as pivot languages that will enable more accurate translations. The dataset is multicentric to enforce representation of multilingual language features. In addition, the dataset goes beyond the sentence level, as it is organized in paragraphs of various lengths. Compared with related machine translation datasets, we show that BOUQuET has a broader representation of domains while simplifying the translation task for non-experts. Therefore, BOUQuET is specially suitable for crowd-source extension for which we are launching a call aiming at collecting a multi-way parallel corpus covering any written language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。