首个面向东南亚多语言AI代理的评估框架,揭示跨语言能力衰减规律
SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages

- 将TauBench适配至五种东南亚语言,构建渐进式本地化评测场景
- 英语代理在仅对话语言变化时表现尚可,全领域本地化后性能大幅下降
- 为多语言AI开发提供可复用的适配流程和诊断工具,适合区域AI研究者
尽管东南亚(SEA)地区人工智能发展迅速,但针对本地语言的智能体能力仍缺乏深入理解,这影响了本土AI的发展。为此,我们提出SEATauBench,首个聚焦于东南亚主权AI的智能体评估框架。该框架将TauBench适配至中文、越南语、泰语、印尼语和菲律宾语五种语言,并在对话语言、工具说明与任务领域逐步本地化的环境中评估多个近期模型。实验发现,当仅改变对话语言时,英语智能体的能力转移尚可;但随着任务上下文进一步本地化,尤其是全领域适应时,性能与鲁棒性显著下降,降幅最大。同时揭示了仅以英语评估难以准确衡量东南亚语言中的智能体能力。更广泛而言,SEATau提供了诊断性基准与可复用的适配流水线,助力构建面向语言多样性区域的可靠多语言智能体。数据与代码可在github.com/SEACrowd/SEATauBench获取。
原文摘要 · Abstract (English)
While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood despite its importance to sovereign AI. To fill this gap, we introduce SEATauBench, the first agent-focused evaluation framework for SEA sovereign AI. SeaTau adapts TauBench to five languages -- Mandarin, Vietnamese, Thai, Indonesian, and Filipino -- and evaluates agents across progressively localized settings that vary the language of user-agent interaction, tool specifications, and task domains. Across three recent models, we find that English agent capabilities transfer reasonably well when only the conversation language changes, but quality and robustness degrade sharply as more task contexts are localized, with the largest losses in full domain adaptation. We also the limits of English-only agent assessment for measuring agent capabilities in SEA languages. More broadly, SeaTau provides a diagnostic benchmark and reusable adaptation pipeline for building reliable multilingual agents for linguistically diverse regions. Data and code can be accessed at github.com/SEACrowd/SEATauBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。