arXiv:2410.00863cs.CL2024-10被引 20

大模型翻译过长影响评估公正性,需调整评价标准。

On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation

  • 分析多个大模型在翻译任务中的冗长输出现象
  • 发现安全、版权和上下文不足是主要诱因
  • 提醒评估时应避免误伤表达更完整的模型

本文研究大模型翻译输出冗长对评估的影响。我们首先在 WMT 2024 机器翻译通用共享任务的多个 LLM 输出中验证了该现象的普遍性。接着识别出冗长的主要触发因素,包括安全顾虑、版权担忧以及短输入查询带来的上下文不足。最后表明,若忽略此行为,将不公平地降低对更冗长模型的评分,无论自动还是人工评估均如此,凸显未来评估需正视这一问题。

原文摘要 · Abstract (English)

This paper investigates the impact of verbose LLM translations on evaluation. We first demonstrate the prevalence of this behavior across several LLM outputs drawn from the WMT 2024 general shared task on machine translation. We then identify the primary triggers of verbosity, including safety, copyright concerns, and insufficient context in short input queries. Finally, we show that ignoring this behavior unfairly penalizes more verbose LLMs according to both automatic and human evaluations, highlighting the need to address this issue for more accurate future evaluations.

机器翻译大模型评估输出冗长

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。