让机器翻译的人工评估像自动评估一样简单高效。
Pearmut: Human Evaluation of Translation Made Trivial
- 轻量级平台,一键部署人工评估流程
- 支持多语言任务与标准评估协议
- 适合需要可靠评估的模型研发团队
人工评估是多语言自然语言处理的黄金标准,但因现有工具工程复杂、操作繁琐,常被自动指标替代。我们提出 Pearmut,一个轻量但功能丰富的平台,使端到端的人工评估如自动评估般便捷。它消除了常见入门障碍,支持多语言任务评估,尤其聚焦机器翻译。平台实现标准评估协议(如 DA、ESA、MQM),可扩展支持新协议。具备文档级上下文、绝对与对比评估、注意力检测、ESAAI 预标注及静态与动态分配策略。使可靠的人工评估成为模型开发与诊断中的常规环节,而非偶尔尝试。
原文摘要 · Abstract (English)
Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is notoriously complex and slow to set up with existing tools with substantial engineering and operational overhead. We introduce Pearmut, a lightweight yet feature-rich platform that makes end-to-end human evaluation as easy to run as automatic evaluation. Pearmut removes common entry barriers and provides support for evaluating multilingual tasks, with a particular focus on machine translation. The platform implements standard evaluation protocols, including DA, ESA, and MQM, and is extensible to support new protocols. It features document-level context, absolute and contrastive evaluation, attention checks, ESAAI pre-annotations and both static and dynamic assignment strategies. Pearmut enables reliable human evaluation to become a practical, routine component of model development and diagnosis rather than an occasional effort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。