arXiv:2601.05101cs.AI2026-01

首个阿拉伯语工具调用基准,揭示语言差异导致性能下降5-10%。

Arabic Prompts with English Tools: A Benchmark

  • 构建首个针对阿拉伯语的LLM工具调用评估基准。
  • 阿拉伯语提示下工具调用准确率平均下降5-10%。
  • 适用于关注多语言AI公平性的研究者与开发者。

大型语言模型(LLMs)已成为众多行业核心推理引擎,通过工具调用执行复杂任务。尽管阿拉伯语原生LLM发展迅速,但评估框架仍滞后,多数现有基准聚焦英语。一个被忽视的关键领域是工具调用:以非英语如阿拉伯语提示时,模型表现尚不清晰,尤其因这些模型通常在大量英语数据上预训练。本文首次提出专用于评估阿拉伯语LLM工具调用与智能体能力的基准。该框架可标准化衡量阿拉伯语智能体工作流的功能准确性和鲁棒性。结果发现显著性能差距:用户使用阿拉伯语交互时,工具调用准确率平均下降5-10%,无论工具描述为阿拉伯语或英语。本基准旨在推动更可靠、语言公平的阿拉伯语AI智能体发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are now integral to numerous industries, increasingly serving as the core reasoning engine for autonomous agents that perform complex tasks through tool-use. While the development of Arabic-native LLMs is accelerating, the benchmarks for evaluating their capabilities lag behind, with most existing frameworks focusing on English. A critical and overlooked area is tool-calling, where the performance of models prompted in non-English languages like Arabic is poorly understood, especially since these models are often pretrained on predominantly English data. This paper addresses this critical gap by introducing the first dedicated benchmark for evaluating the tool-calling and agentic capabilities of LLMs in the Arabic language. Our work provides a standardized framework to measure the functional accuracy and robustness of models in Arabic agentic workflows. Our findings reveal a huge performance gap: when users interact in Arabic, tool-calling accuracy drops by an average of 5-10\%, regardless of whether the tool descriptions themselves are in Arabic or English. By shedding light on these critical challenges, this benchmark aims to foster the development of more reliable and linguistically equitable AI agents for Arabic-speaking users.

大模型多语言工具调用阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。