arXiv:2601.07020cs.CLcs.AI2026-01中稿 · EACL 2026 SIGTURK被引 4

构建首个土耳其语大模型评估基准,覆盖21类任务。

TurkBench: A Benchmark for Evaluating Turkish Large Language Models

  • 设计包含8151样本的土耳其语评测集,分6大类21子任务。
  • 涵盖知识、语法、推理等能力,数据具文化相关性。
  • 适合研究土耳其语AI或开发多语言模型的团队使用。

随着大语言模型的快速发展,针对特定语言的全面评估基准需求日益迫切。尽管英语模型评估已取得显著进展,但具有独特语言特征的土耳其语等领域仍缺乏完善评测体系。本文提出TurkBench,一个专为评估土耳其语生成式大模型而设计的综合性基准,包含8,151个数据样本,覆盖21个不同子任务,分为六大评估类别:知识、语言理解、推理、内容审核、土耳其语语法与词汇、指令遵循。多样化的任务设计和具有文化相关性的数据,为研究人员和开发者提供评估模型表现及定位改进方向的重要工具。该基准已发布于Hugging Face平台,支持在线提交评测。

原文摘要 · Abstract (English)

With the recent surge in the development of large language models, the need for comprehensive and language-specific evaluation benchmarks has become critical. While significant progress has been made in evaluating English-language models, benchmarks for other languages, particularly those with unique linguistic characteristics such as Turkish, remain less developed. Our study introduces TurkBench, a comprehensive benchmark designed to assess the capabilities of generative large language models in the Turkish language. TurkBench involves 8,151 data samples across 21 distinct subtasks. These are organized under six main categories of evaluation: Knowledge, Language Understanding, Reasoning, Content Moderation, Turkish Grammar and Vocabulary, and Instruction Following. The diverse range of tasks and the culturally relevant data would provide researchers and developers with a valuable tool for evaluating their models and identifying areas for improvement. We further publish our benchmark for online submissions at https://huggingface.co/turkbench

语言模型评测基准土耳其语多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。