arXiv:2504.20371cs.CL2025-04EMNLP被引 5

评测大模型在多领域翻译中的消歧能力,发现其表现仍有巨大提升空间。

DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translation

  • 构建带领域标注的多领域翻译测试集,精准刻画词汇歧义问题。
  • 在4种语言对、13个领域上测试主流大模型,发现其消歧准确率普遍不足60%。
  • 提供可复用的提示策略与评估指标,适合关注翻译质量与模型鲁棒性的研究者。

当前大型语言模型(LLMs)在机器翻译中取得了显著成果,但在多领域翻译(MDT)中的表现仍不理想。由于词语在不同领域含义差异显著,导致MDT存在显著歧义。因此,评估LLMs在多领域翻译中的消歧能力仍是开放性问题。为此,本文提出DMDTEval——一个系统性的评估框架,包含三个核心部分:(1) 构建带有领域歧义标注的翻译测试集;(2) 整合多样化的消歧提示策略;(3) 设计精确的消歧评估指标,并研究多种提示策略在多个前沿大模型上的效果。我们在4种语言对和13个领域上进行全面实验,结果揭示了若干关键发现,有望为提升大模型消歧能力的研究提供重要参考。

原文摘要 · Abstract (English)

Currently, Large Language Models (LLMs) have achieved remarkable results in machine translation. However, their performance in multi-domain translation (MDT) is less satisfactory, the meanings of words can vary across different domains, highlighting the significant ambiguity inherent in MDT. Therefore, evaluating the disambiguation ability of LLMs in MDT, remains an open problem. To this end, we present an evaluation and analysis of LLMs on disambiguation in multi-domain translation (DMDTEval), our systematic evaluation framework consisting of three critical aspects: (1) we construct a translation test set with multi-domain ambiguous word annotation, (2) we curate a diverse set of disambiguation prompt strategies, and (3) we design precise disambiguation metrics, and study the efficacy of various prompt strategies on multiple state-of-the-art LLMs. We conduct comprehensive experiments across 4 language pairs and 13 domains, our extensive experiments reveal a number of crucial findings that we believe will pave the way and also facilitate further research in the critical area of improving the disambiguation of LLMs.

多领域翻译大模型评估歧义消解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。