arXiv:2508.16431cs.CLcs.AI2025-08被引 7

首个全面评估土耳其语大模型语言理解与文化能力的基准测试

Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish

  • 构建涵盖23项任务的统一评测框架,覆盖语法纠错、翻译等
  • 33个开源模型对比显示土耳其专属模型表现不及通用模型
  • 聚焦文化语境任务,适合研究土耳其语AI或跨语言模型评估者

我们提出Cetvel,一个面向土耳其语大语言模型的综合性评估基准。现有土耳其语评测数据集普遍存在任务多样性不足或缺乏文化相关性的问题。Cetvel通过整合23项任务(分属7个类别),包括基于土耳其历史和习语的问答、语法纠错和机器翻译等,全面反映土耳其语的语言与文化特征。我们评估了33个开源模型(参数量达70B),涵盖不同模型家族与指令微调范式。实验表明,尽管专为土耳其语设计,但土耳其定制化指令微调模型整体表现仍逊于多语言或通用模型(如Llama 3和Mistral)。此外,语法纠错和抽取式问答任务在区分模型能力方面尤为敏感。Cetvel为提升土耳其语大模型的发展与评估提供了全面且文化契合的评测工具。

原文摘要 · Abstract (English)

We introduce Cetvel, a comprehensive benchmark designed to evaluate large language models (LLMs) in Turkish. Existing Turkish benchmarks often lack either task diversity or culturally relevant content, or both. Cetvel addresses these gaps by combining a broad range of both discriminative and generative tasks ensuring content that reflects the linguistic and cultural richness of Turkish language. Cetvel covers 23 tasks grouped into seven categories, including tasks such as grammatical error correction, machine translation, and question answering rooted in Turkish history and idiomatic language. We evaluate 33 open-weight LLMs (up to 70B parameters) covering different model families and instruction paradigms. Our experiments reveal that Turkish-centric instruction-tuned models generally underperform relative to multilingual or general-purpose models (e.g. Llama 3 and Mistral), despite being tailored for the language. Moreover, we show that tasks such as grammatical error correction and extractive question answering are particularly discriminative in differentiating model capabilities. Cetvel offers a comprehensive and culturally grounded evaluation suite for advancing the development and assessment of LLMs in Turkish.

语言模型土耳其语评估基准文化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。