arXiv:2501.00593cs.CL2025-01被引 11

构建土耳其语大模型评估标准,填补资源有限语言评测空白

Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation

  • 基于62个学科6200道题构建土耳其语评测基准TR-MMLU
  • 首次系统评估主流大模型在土耳其语任务中的表现与局限
  • 适合关注低资源语言、教育文本理解的研究者使用

大语言模型在理解和生成自然语言方面取得显著进展,但在资源有限语言如土耳其语上的评估仍具挑战。为此,本文提出土耳其MMLU(TR-MMLU)基准,一个涵盖62个领域、6200道多项选择题的综合性评估框架,源自28万道题、覆盖67个学科和800多个主题的土耳其教育体系数据集。该基准提供透明、可复现且文化相关的评测工具,可作为土耳其NLP研究的标准框架,支持对模型处理土耳其语能力的深入分析,并推动更鲁棒、准确的语言模型发展。本研究对前沿大模型在TR-MMLU上的表现进行评估,揭示了分词方式与微调策略的影响,指出了模型设计的改进方向。

原文摘要 · Abstract (English)

Language models have made remarkable advancements in understanding and generating human language, achieving notable success across a wide array of applications. However, evaluating these models remains a significant challenge, particularly for resource-limited languages such as Turkish. To address this gap, we introduce the Turkish MMLU (TR-MMLU) benchmark, a comprehensive evaluation framework designed to assess the linguistic and conceptual capabilities of large language models (LLMs) in Turkish. TR-MMLU is constructed from a carefully curated dataset comprising 6200 multiple-choice questions across 62 sections, selected from a pool of 280000 questions spanning 67 disciplines and over 800 topics within the Turkish education system. This benchmark provides a transparent, reproducible, and culturally relevant tool for evaluating model performance. It serves as a standard framework for Turkish NLP research, enabling detailed analyses of LLMs' capabilities in processing Turkish text and fostering the development of more robust and accurate language models. In this study, we evaluate state-of-the-art LLMs on TR-MMLU, providing insights into their strengths and limitations for Turkish-specific tasks. Our findings reveal critical challenges, such as the impact of tokenization and fine-tuning strategies, and highlight areas for improvement in model design. By setting a new standard for evaluating Turkish language models, TR-MMLU aims to inspire future innovations and support the advancement of Turkish NLP research.

大模型评估土耳其语多任务评测低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。