测试机器翻译如何改变文本复杂度,发现高难度文本更难译且译文复杂度会变化。
ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation

- 用CEFR等级衡量文本复杂度,评估翻译对复杂度的影响
- 高CEFR文本更难翻译,多数语言译文复杂度与原文不同
- 适合研究多语种教育内容生成和翻译难度评估的学者
当文本被翻译时,其复杂度是否得以保留?我们提出ComplexityMT,一个评估文本复杂度与机器翻译相互作用的新基准,以欧洲共同语言参考框架(CEFR)等级作为文本复杂度的衡量标准。在阿拉伯语、荷兰语、英语、法语、印地语和俄语六种语言上,我们评估了三个开源模型、一个闭源模型及一个商用机器翻译系统在两项任务上的表现:一、CEFR等级与翻译难度的相关性;二、源文本CEFR等级在翻译后的变化。实验表明,更高CEFR等级的文本更难翻译,且大多数语言的机器翻译会导致目标文本的CEFR等级相较于源文本发生偏移。这些发现为多语种教学内容生成与翻译难度估计的研究者提供了新视角。
原文摘要 · Abstract (English)
When a text is translated, does the translation retain the complexity of the original? We introduce ComplexityMT, a new challenge for assessing how text complexity and machine translation interact with and influence each other, using the Common European Framework of Reference for Languages (CEFR) levels as the measure of text complexity. Across six languages, including Arabic, Dutch, English, French, Hindi, and Russian, we evaluate three open-weight models, one closed model, and a commercial machine translation system on two tasks: i) correlation of CEFR with translation difficulty, and ii) shifts in CEFR levels of the source texts. Our experiments show that higher CEFR levels make texts more difficult to translate, and that machine translation shifts the CEFR level of the target text compared to the original source, for most languages. These findings provide new insights for researchers and practitioners working on multilingual pedagogical content generation and machine translation difficulty estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。