arXiv:2608.08775cs.CL2026-08

测试多语言AI代理性能,发现英文能力无法直接迁移到其他语言。

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

论文配图:OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents
图 1 · 摘自论文原文
  • 用机器翻译扩展英文基准至10种语言,结合人工校准验证器
  • 所有代理在非英语语言上平均下降8.8-18.4分(pass@3),工具调度是主要短板
  • 模型本身是差距主因,非拉丁字母语言受语素线索丢失影响显著

智能体评估基准旨在衡量AI代理在真实多工具环境中规划、搜索、执行和恢复的能力,但几乎全部为英文。随着AI代理面向全球多语言用户部署,其在英文中表现的智能体能力是否可迁移至其他语言仍未知。我们提出OmnilingualGAIA2,即对GAIA2基准的机器翻译扩展(部分经人工专家验证),覆盖十种目标语言,涵盖五种书写系统,并配备本地化的人工校准多语言验证器。评估七种前沿及开源模型,发现普遍存在8.8–18.4的跨语言差距(pass@3),该差距在不同模型间不对称,集中于工具编排而非定量推理,且不随模型规模缩小。分层错误归因显示,55%的差距源于模型自身,翻译污染仅占6.4%场景-语言组合。人工语言学分析进一步揭示:非拉丁文字语言中形态线索丢失与歧义放大是主要失败机制。结果表明,多语言智能体评估应成为全球部署代理的标准报告流程。

原文摘要 · Abstract (English)

Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question. We introduce OmnilingualGAIA2, a machine-translated expansion (with partial human- expert validation) of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier. Evaluating seven frontier and open-weight agents, we find a universal cross-lingual gap of 8.8-18.4 pass@3 points that is agent-asymmetric in magnitude, concentrates on tool-orchestration rather than quantitative reasoning, and does not close with model scale. A stratified error attribution decomposes the gap as predominantly model-driven (55%), with a bounded translation-contamination floor of only 6.4% of scenario-language pairs. Human-expert linguistic analysis further identifies morphological cue loss and amplified ambiguity as the primary failure mechanisms in non-Latin-script languages. Our results argue that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents.

多语言智能体评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。