arXiv:2502.06298cs.CLcs.AI2025-02NAACL被引 19

用东南亚真实语言任务评测大模型,更准

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

  • 基于真实东南亚考试与日常对话构建评测集
  • 在本地化任务上,新基准比翻译数据更有效
  • 适合评估多语言大模型在区域场景的表现

本研究推出两个新基准——SeaExam 和 SeaBench,用于评估大语言模型(LLMs)在东南亚(SEA)应用环境中的能力。与现有主要依赖英文翻译的多语言数据集不同,这两个基准均基于东南亚地区的实际应用场景构建。SeaExam 源自区域教育考试,涵盖地方历史、文学等学科;SeaBench 则围绕多轮开放式任务设计,反映东南亚社区的日常互动。实验表明,相比翻译生成的基准,SeaExam 与 SeaBench 能更有效地区分 LLM 在东南亚语言任务上的表现,凸显使用真实查询评估多语言模型能力的重要性。

原文摘要 · Abstract (English)

This study introduces two novel benchmarks, SeaExam and SeaBench, designed to evaluate the capabilities of Large Language Models (LLMs) in Southeast Asian (SEA) application scenarios. Unlike existing multilingual datasets primarily derived from English translations, these benchmarks are constructed based on real-world scenarios from SEA regions. SeaExam draws from regional educational exams to form a comprehensive dataset that encompasses subjects such as local history and literature. In contrast, SeaBench is crafted around multi-turn, open-ended tasks that reflect daily interactions within SEA communities. Our evaluations demonstrate that SeaExam and SeaBench more effectively discern LLM performance on SEA language tasks compared to their translated benchmarks. This highlights the importance of using real-world queries to assess the multilingual capabilities of LLMs.

多语言评测大模型东南亚

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。