arXiv:2602.22207cs.CLcs.AI2026-02ACL被引 1

自动化翻译框架提升多语言评测数据质量,避免语义丢失。

Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets

  • 用多轮排名T-RANK和自改进策略提升翻译准确性
  • 在8种东欧南欧语言中实现高质量评测集翻译
  • 适合需要可靠多语言模型评估的研究者使用

当前多语言大模型评估的可靠性因翻译评测集质量参差而受损。现有资源常出现语义漂移和上下文丢失,导致性能指标失真。本文提出一个全自动化框架,可规模化生成高质量数据集与评测集。通过引入测试时计算扩展策略,包括通用自改进(USI)和提出的多轮排名方法T-RANK,显著提升输出质量。该框架确保本地化过程中保留原任务结构与语言细微差别。我们将此方法应用于将主流评测集翻译成乌克兰语、保加利亚语、斯洛伐克语、罗马尼亚语、立陶宛语、爱沙尼亚语、土耳其语和希腊语共8种语言。基于参考标准和大模型评分的评估表明,我们的翻译结果优于现有资源,使下游模型评估更准确。我们已开源该框架及优化后的评测集,以推动稳健、可复现的多语言人工智能发展。

原文摘要 · Abstract (English)

The reliability of multilingual Large Language Model (LLM) evaluation is currently compromised by the inconsistent quality of translated benchmarks. Existing resources often suffer from semantic drift and context loss, which can lead to misleading performance metrics. In this work, we present a fully automated framework designed to address these challenges by enabling scalable, high-quality translation of datasets and benchmarks. We demonstrate that adapting test-time compute scaling strategies, specifically Universal Self-Improvement (USI) and our proposed multi-round ranking method, T-RANK, allows for significantly higher quality outputs compared to traditional pipelines. Our framework ensures that benchmarks preserve their original task structure and linguistic nuances during localization. We apply this approach to translate popular benchmarks and datasets into eight Eastern and Southern European languages (Ukrainian, Bulgarian, Slovak, Romanian, Lithuanian, Estonian, Turkish, Greek). Evaluations using both reference-based metrics and LLM-as-a-judge show that our translations surpass existing resources, resulting in more accurate downstream model assessment. We release both the framework and the improved benchmarks to facilitate robust and reproducible multilingual AI development.

多语言评估自动翻译评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。