arXiv:2604.24929cs.CLcs.AI2026-04被引 1

重构多语言智能体评测流程,提升跨语言任务一致性

GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation

  • 通过人工校验与自动化检查,确保任务功能与文化语境对齐
  • 多语言代理成功率最高提升32.7%,最接近英文表现仅差3.1%
  • 适合关注跨语言AI评估公平性的研究者与开发者

当前智能体评测仍以英语为主,多数多语言版本依赖机器翻译与有限后编辑,易导致查询-回答错位或文化偏差。本文提出一种新流程,通过显式功能对齐、文化适配和难度校准,结合自动检测与人工审核,重构多语言评测体系。基于该流程,我们推出了GAIA-v2-LILT,涵盖五种非英语语言的重新审计版本。实验表明,该方法相较最小化翻译版本,代理成功率最高提升32.7%,最优情况下与英语性能差距缩小至3.1%以内,但多数语言仍存在显著差距。这表明多语言性能差异中相当部分源于评测本身带来的测量误差,强调在跨语言迁移时需进行任务级对齐。数据已作为MAPS包发布于Hugging Face,代码开源在GitHub。

原文摘要 · Abstract (English)

Agent benchmarks remain largely English-centric, while their multilingual versions are often built with machine translation (MT) and limited post-editing. We argue that, for agentic tasks, this minimal workflow can easily break benchmark validity through query-answer misalignment or culturally off-target context. We propose a refined workflow for adapting English benchmarks into multiple languages with explicit functional alignment, cultural alignment, and difficulty calibration using both automated checks and human review. Using this workflow, we introduce GAIA-v2-LILT, a re-audited multilingual extension of GAIA covering five non-English languages. In experiments, our workflow improves agent success rates by up to 32.7% over minimally translated versions, bringing the closest audited setting to within 3.1% of English performance while substantial gaps remain in many other cases. This indicates that a substantial share of the multilingual performance gap is benchmark-induced measurement error, motivating task-level alignment when adapting English benchmarks across languages. The data is available as part of the MAPS package at https://huggingface.co/datasets/Fujitsu-FRE/MAPS/viewer/GAIA-v2-LILT. We also release the code used in our experiments at https://github.com/lilt/gaia-v2-lilt.

智能体评测多语言文化对齐基准校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。