arXiv:2505.15935cs.DBcs.CL2025-05Conference of the …被引 6

首个多语言智能体评估基准,揭示非英语下性能与安全下降问题

MAPS: A Multilingual Benchmark for Agent Performance and Security

  • 基于四大主流基准翻译成11种语言,构建805个任务9660个实例的多语言评估集
  • 实测显示非英语环境下智能体表现与安全性普遍下降,降幅与翻译输入量正相关
  • 为多语言智能体公平性研究提供标准化框架,适合关注AI可及性与安全的研究者

基于大语言模型的智能体系统在工具调用与记忆交互方面快速进步,但现有模型在多语言场景中表现不佳,导致性能降低与安全风险上升。这引发对非英语用户使用体验的担忧。尽管已有初步研究探索智能体评估与多语言交互,但尚无涵盖多领域、具备安全意识的综合性多语言评估基准。为此,本文提出MAPS,一个覆盖十一种语言的多语言智能体评估基准,基于GAIA(真实任务)、SWE-Bench(代码生成)、MATH(数学推理)和Agent Security Benchmark(安全测试)四个主流基准进行翻译,共生成805个独特任务与9,660个语言特异性实例。实验证明,从英语转至其他语言时,智能体在性能与安全性上均出现显著退化,且退化程度与翻译输入量相关。本工作首次建立多语言智能体的标准化评估框架,推动更公平、可靠、可访问的智能体发展。MAPS已公开于Hugging Face:https://huggingface.co/datasets/Fujitsu-FRE/MAPS

原文摘要 · Abstract (English)

Agentic AI systems, which build on Large Language Models (LLMs) and interact with tools and memory, have rapidly advanced in capability and scope. Yet, since LLMs have been shown to struggle in multilingual settings, typically resulting in lower performance and reduced safety, agentic systems risk inheriting these limitations. This raises concerns about the accessibility of such systems, as users interacting in languages other than English may encounter unreliable or security-critical agent behavior. Despite growing interest in evaluating agentic AI and recent initial efforts toward multilingual interaction, existing benchmarks do not yet provide a comprehensive, multi-domain, security-aware evaluation of multilingual agentic systems. To address this gap, we propose MAPS, a multilingual benchmark suite designed to evaluate agentic AI systems across diverse languages and tasks. MAPS builds on four widely used agentic benchmarks - GAIA (real-world tasks), SWE-Bench (code generation), MATH (mathematical reasoning), and the Agent Security Benchmark (security). We translate each dataset into eleven diverse languages, resulting in 805 unique tasks and 9,660 total language-specific instances - enabling a systematic analysis of the Multilingual Effect on AI agents' performance and robustness. Empirically, we observe a degradation in both performance and security when transitioning from English to other languages, with severity varying by task and correlating with the amount of translated input. This work establishes the first standardized evaluation framework for multilingual agentic AI, encouraging future research towards equitable, reliable, and accessible agentic AI. MAPS benchmark suite is publicly available at https://huggingface.co/datasets/Fujitsu-FRE/MAPS

多语言智能体评估基准安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。