构建多语言赛事购票评估基准,揭示大模型跨语言能力差距
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
- 基于六种语言的足球票务场景,模拟真实地域文化差异
- 主流模型在多语言任务中准确率最高达87.3%,但跨语言表现波动超20%
- 适合关注大模型全球化部署与文化适配的研究者
大型语言模型正越来越多地作为任务导向型智能体部署,其成功依赖于在真实多语言环境下生成准确函数调用的能力。然而现有代理评估普遍忽视文化和语言多样性,多依赖单语或简单翻译的基准。本文提出Ticket-Bench,一个面向任务型场景的多语言代理评估基准,模拟六种主要语言(葡萄牙语、英语、西班牙语、德语、意大利语、法语)下的足球票务购买场景。通过引入本地化球队、城市及用户画像,提升真实性。评估涵盖多种商用和开源大模型,测量其函数调用准确率与跨语言一致性。结果表明,推理导向模型(如GPT-5、Qwen3-235B)表现领先,但跨语言性能差异显著,最高达21.4个百分点。研究强调需发展具备文化敏感性的多语言评估基准,以推动鲁棒性智能体的发展。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as task-oriented agents, where success depends on their ability to generate accurate function calls under realistic, multilingual conditions. However, existing agent evaluations largely overlook cultural and linguistic diversity, often relying on monolingual or naively translated benchmarks. We introduce Ticket-Bench, a benchmark for multilingual agent evaluation in task-oriented scenarios. Ticket-Bench simulates the domain of soccer ticket purchases across six major languages: Portuguese, English, Spanish, German, Italian, and French. Using localized teams, cities, and user profiles to provide a higher level of realism. We evaluate a wide range of commercial and open-source LLMs, measuring function-calling accuracy and consistency across languages. Results show that reasoning-oriented models (e.g., GPT-5, Qwen3-235B) dominate performance but still exhibit notable cross-lingual disparities. These findings underscore the need for culturally aware, multilingual benchmarks to guide the development of robust LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。