arXiv:2412.11615cs.CL2024-12NAACL被引 1

一站式工具,全面评估机器翻译质量与偏见

MT-LENS: An all-in-one Toolkit for Better Machine Translation Evaluation

  • 扩展LM-eval-harness,支持多种翻译评估任务
  • 覆盖质量、性别偏见、毒性、拼写鲁棒性等维度
  • 可视化对比,适合研究者与工程师快速分析模型

我们提出MT-LENS,一个面向机器翻译(MT)系统多维度评估的框架,涵盖翻译质量、性别偏见检测、毒性内容识别及对拼写错误的鲁棒性。尽管已有多个工具广泛用于大语言模型(LLM)能力评测,但现有评估工具难以全面衡量MT性能的多样性。MT-LENS通过扩展LM-eval-harness,支持当前最先进的数据集和多种评估指标,提供用户友好的平台,实现系统间对比与交互式翻译分析。该框架旨在推动超越传统翻译质量评价的评估策略普及,帮助研究人员和工程师更全面理解神经机器翻译(NMT)模型表现,并便捷测量系统偏差。

原文摘要 · Abstract (English)

We introduce MT-LENS, a framework designed to evaluate Machine Translation (MT) systems across a variety of tasks, including translation quality, gender bias detection, added toxicity, and robustness to misspellings. While several toolkits have become very popular for benchmarking the capabilities of Large Language Models (LLMs), existing evaluation tools often lack the ability to thoroughly assess the diverse aspects of MT performance. MT-LENS addresses these limitations by extending the capabilities of LM-eval-harness for MT, supporting state-of-the-art datasets and a wide range of evaluation metrics. It also offers a user-friendly platform to compare systems and analyze translations with interactive visualizations. MT-LENS aims to broaden access to evaluation strategies that go beyond traditional translation quality evaluation, enabling researchers and engineers to better understand the performance of a NMT model and also easily measure system's biases.

机器翻译评估工具偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。