arXiv:2510.17460cs.CL2025-10

首个乌尔都习语翻译数据集,评测大模型跨语言文化表达能力

Evaluating Large Language Models on Urdu Idiom Translation

  • 构建首个乌尔都语到英语习语翻译数据集,支持原生与罗马字母双书写形式
  • 提示工程提升习语翻译效果,原生乌尔都文本比罗马转写更准确
  • 适合低资源语言翻译、跨文化语义理解研究者参考

习语翻译仍是机器翻译中的重大挑战,尤其对乌尔都语等低资源语言关注不足。为推动该领域研究,我们首次构建了乌尔都语至英语习语翻译的评估数据集,涵盖原生乌尔都语与罗马乌尔都语两种书写形式,并标注了标准英文对应译文。我们在该任务上评估多个开源大语言模型(LLMs)和神经机器翻译(NMT)系统,重点考察其保留习语与文化内涵的能力。采用BLEU、BERTScore、COMET和XCOMET等自动指标评估翻译质量。结果表明,提示工程相比直接翻译显著提升习语翻译表现,但不同提示类型间差异较小;跨书写形式对比显示,文本表征对翻译质量影响显著,原生乌尔都语输入产生的习语翻译更准确。

原文摘要 · Abstract (English)

Idiomatic translation remains a significant challenge in machine translation, especially for low resource languages such as Urdu, and has received limited prior attention. To advance research in this area, we introduce the first evaluation datasets for Urdu to English idiomatic translation, covering both Native Urdu and Roman Urdu scripts and annotated with gold-standard English equivalents. We evaluate multiple open-source Large Language Models (LLMs) and Neural Machine Translation (NMT) systems on this task, focusing on their ability to preserve idiomatic and cultural meaning. Automatic metrics including BLEU, BERTScore, COMET, and XCOMET are used to assess translation quality. Our findings indicate that prompt engineering enhances idiomatic translation compared to direct translation, though performance differences among prompt types are relatively minor. Moreover, cross script comparisons reveal that text representation substantially affects translation quality, with Native Urdu inputs producing more accurate idiomatic translations than Roman Urdu.

习语翻译低资源语言大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。