arXiv:2506.11467cs.CLcs.SI2025-06

为低资源语言翻译系统设计游戏化评估与招募平台

A Gamified Evaluation and Recruitment Platform for Low Resource Language Machine Translation Systems

  • 构建游戏化平台吸引并招募低资源语言母语者参与评估
  • 解决低资源语言数据集和人工评估者双重短缺问题
  • 适合语言资源匮乏的NLP研究者及翻译系统开发者

人类评估者在大语言模型评测中至关重要。对于低资源语言(LRLs)的机器翻译(MT)系统,这一需求尤为突出,因为主流自动化指标多为基于字符串的,难以捕捉系统的语义细微差别。具备语言专长的人类评估者可有效评测翻译的准确性、流畅性等关键维度。然而,低资源语言面临数据集和评估者双重匮乏的困境。本文首先系统回顾现有评估流程,提出一个旨在缓解数据与人力资源缺口的平台设计方案:一个面向低资源语言机器翻译开发者的招募与游戏化评估平台。同时讨论了该平台的评估挑战及其在更广泛自然语言处理研究中的潜在应用。

原文摘要 · Abstract (English)

Human evaluators provide necessary contributions in evaluating large language models. In the context of Machine Translation (MT) systems for low-resource languages (LRLs), this is made even more apparent since popular automated metrics tend to be string-based, and therefore do not provide a full picture of the nuances of the behavior of the system. Human evaluators, when equipped with the necessary expertise of the language, will be able to test for adequacy, fluency, and other important metrics. However, the low resource nature of the language means that both datasets and evaluators are in short supply. This presents the following conundrum: How can developers of MT systems for these LRLs find adequate human evaluators and datasets? This paper first presents a comprehensive review of existing evaluation procedures, with the objective of producing a design proposal for a platform that addresses the resource gap in terms of datasets and evaluators in developing MT systems. The result is a design for a recruitment and gamified evaluation platform for developers of MT systems. Challenges are also discussed in terms of evaluating this platform, as well as its possible applications in the wider scope of Natural Language Processing (NLP) research.

机器翻译低资源语言评估平台游戏化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。