arXiv:2512.11074cs.CLcs.AI2025-12被引 4

扩展多语言图文翻译数据集,支持中、阿、西、俄等五种语言和双字形中文。

MultiScript30k: Leveraging Multilingual Embeddings to Extend Cross Script Parallel Data

  • 用NLLB大模型批量翻译英文图文数据至多种语言和文字。
  • 生成超3万句跨语言平行语料,支持拉丁、阿拉伯、汉字等五种书写系统。
  • 适合做多语言图文翻译、跨文化通用性研究的学者使用。

Multi30k是多模态机器翻译(MMT)领域常用的数据集,仅包含捷克语、英语、法语和德语四种欧洲拉丁文语言,限制了其他语言的研究。为突破此局限,本文提出MultiScript30k,通过NLLB200-3.3B模型将Multi30k英文版(Multi30k-En)翻译为阿拉伯语(Ar)、西班牙语(Es)、乌克兰语(Uk)、简体中文(Zh_Hans)和繁体中文(Zh_Hant),共生成超过30,000句平行文本。相似性分析显示,除繁体中文外,所有语言翻译结果与原数据的余弦相似度均高于0.8,对称KL散度低于0.000251。COMETKiwi评估显示,其翻译质量与已有成果相当,其中乌语版本比现有版本高6.4%(每分割段),阿语版本接近原有基准,但繁体中文表现与之前扩展数据相近。

原文摘要 · Abstract (English)

Multi30k is frequently cited in the multimodal machine translation (MMT) literature, offering parallel text data for training and fine-tuning deep learning models. However, it is limited to four languages: Czech, English, French, and German. This restriction has led many researchers to focus their investigations only on these languages. As a result, MMT research on diverse languages has been stalled because the official Multi30k dataset only represents European languages in Latin scripts. Previous efforts to extend Multi30k exist, but the list of supported languages, represented language families, and scripts is still very short. To address these issues, we propose MultiScript30k, a new Multi30k dataset extension for global languages in various scripts, created by translating the English version of Multi30k (Multi30k-En) using NLLB200-3.3B. The dataset consists of over \(30000\) sentences and provides translations of all sentences in Multi30k-En into Ar, Es, Uk, Zh\_Hans and Zh\_Hant. Similarity analysis shows that Multi30k extension consistently achieves greater than \(0.8\) cosine similarity and symmetric KL divergence less than \(0.000251\) for all languages supported except Zh\_Hant which is comparable to the previous Multi30k extensions ArEnMulti30k and Multi30k-Uk. COMETKiwi scores reveal mixed assessments of MultiScript30k as a translation of Multi30k-En in comparison to the related work. ArEnMulti30k scores nearly equal MultiScript30k-Ar, but Multi30k-Uk scores $6.4\%$ greater than MultiScript30k-Uk per split.

多语言图文翻译数据集扩展NLLB

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。