arXiv:2501.04473cs.CL2025-01被引 10

无需参考文本,用大模型评估低资源语言翻译质量

When LLMs Struggle: Reference-less Translation Evaluation for Low-resource Languages

  • 用提示工程与指令微调提升LLM在无参考翻译下的评估能力
  • 编码器微调模型在零/少样本下表现优于提示方法
  • 发现分词、音译与专有名词是主要错误来源

本文研究低资源语言对的无参考机器翻译质量评估(QE),即在没有标准参考译文的情况下,为翻译结果给出0-100的质量评分。该任务属于跨语言理解挑战,需在零样本或少样本场景下评估大语言模型(LLMs)性能。研究通过基于标注指南的新颖提示设计进行指令微调,并全面评估其效果。结果表明,基于提示的方法表现不如编码器架构微调后的QE模型。误差分析揭示了分词不当、音译错误及专有名词处理问题,提示需优化大模型在跨语言任务中的预训练策略。相关数据与模型已公开,供后续研究使用。

原文摘要 · Abstract (English)

This paper investigates the reference-less evaluation of machine translation for low-resource language pairs, known as quality estimation (QE). Segment-level QE is a challenging cross-lingual language understanding task that provides a quality score (0-100) to the translated output. We comprehensively evaluate large language models (LLMs) in zero/few-shot scenarios and perform instruction fine-tuning using a novel prompt based on annotation guidelines. Our results indicate that prompt-based approaches are outperformed by the encoder-based fine-tuned QE models. Our error analysis reveals tokenization issues, along with errors due to transliteration and named entities, and argues for refinement in LLM pre-training for cross-lingual tasks. We release the data, and models trained publicly for further research.

机器翻译质量评估低资源语言LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。