arXiv:2410.06338cs.CL2024-10被引 2

LLM能准确评估UGC机器翻译质量,但存在回复拒绝和输出不稳定问题。

Are Large Language Models State-of-the-art Quality Estimators for Machine Translation of User-generated Content?

  • 用参数高效微调(PEFT)提升LLM对UGC翻译质量的预测能力
  • 在无参考译文情况下,PEFT-LMM得分预测准确率优于基线模型
  • 适合关注翻译质量评估且需可解释性的研究者使用

本文研究大型语言模型(LLMs)是否可作为无需参考译文的用户生成内容(UGC)机器翻译质量评估的最先进工具,尤其针对包含情感表达的文本。我们基于已有情感相关数据集,结合人工标注错误,采用多维质量度量(Multi-dimensional Quality Metrics)计算评估分数。在上下文学习与参数高效微调(PEFT)两种场景下,对比多种LLM与微调基线模型的表现。结果表明,对LLM进行PEFT后,在评分预测上表现更优,且能提供人类可理解的解释。然而,人工分析发现其在评估过程中仍存在拒绝响应提示、输出不稳定等问题。

原文摘要 · Abstract (English)

This paper investigates whether large language models (LLMs) are state-of-the-art quality estimators for machine translation of user-generated content (UGC) that contains emotional expressions, without the use of reference translations. To achieve this, we employ an existing emotion-related dataset with human-annotated errors and calculate quality evaluation scores based on the Multi-dimensional Quality Metrics. We compare the accuracy of several LLMs with that of our fine-tuned baseline models, under in-context learning and parameter-efficient fine-tuning (PEFT) scenarios. We find that PEFT of LLMs leads to better performance in score prediction with human interpretable explanations than fine-tuned models. However, a manual analysis of LLM outputs reveals that they still have problems such as refusal to reply to a prompt and unstable output while evaluating machine translation of UGC.

机器翻译LLM评估UGCPEFT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。