arXiv:2508.13044cs.CL2025-08被引 2

首个专为土耳其语大模型设计的综合性评测基准。

Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları

  • 基于62个教育领域6200道选择题构建评测集
  • 揭示主流大模型在土耳其语理解上的短板
  • 助力土耳其语NLP研究与模型优化

大语言模型在理解和生成人类语言方面取得了显著进展,并在各类应用中表现优异。然而,对这些模型的评估仍具挑战性,尤其是在资源有限的语言如土耳其语中。为解决此问题,我们提出了土耳其语MMLU(TR-MMLU)评测基准,这是一个全面的评估框架,旨在衡量大语言模型(LLMs)在土耳其语中的语言与概念能力。TR-MMLU基于精心整理的数据集,涵盖62个教育领域的6200道多项选择题。该基准为土耳其语自然语言处理研究提供了标准评估体系,使研究人员能够深入分析大模型处理土耳其语文本的能力。本研究对前沿大模型在TR-MMLU上的表现进行了评估,揭示了模型设计中的改进空间。TR-MMLU为推动土耳其语NLP研究发展树立了新标准,并激励未来创新。

原文摘要 · Abstract (English)

Language models have made significant advancements in understanding and generating human language, achieving remarkable success in various applications. However, evaluating these models remains a challenge, particularly for resource-limited languages like Turkish. To address this issue, we introduce the Turkish MMLU (TR-MMLU) benchmark, a comprehensive evaluation framework designed to assess the linguistic and conceptual capabilities of large language models (LLMs) in Turkish. TR-MMLU is based on a meticulously curated dataset comprising 6,200 multiple-choice questions across 62 sections within the Turkish education system. This benchmark provides a standard framework for Turkish NLP research, enabling detailed analyses of LLMs' capabilities in processing Turkish text. In this study, we evaluated state-of-the-art LLMs on TR-MMLU, highlighting areas for improvement in model design. TR-MMLU sets a new standard for advancing Turkish NLP research and inspiring future innovations.

土耳其语大模型评测语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。