arXiv:2512.22100cs.CLcs.AI2025-12被引 2

首个土耳其语通用语言理解与情感分析基准,助力NLP模型评测

Introducing TrGLUE and SentiTurca: A Comprehensive Benchmark for Turkish General Language Understanding and Sentiment Analysis

  • 构建TrGLUE和SentiTurca双基准,覆盖土耳其语多任务NLU与情感分析
  • 采用半自动化标注流程,结合大模型与人工校验,保证数据质量与自然性
  • 提供完整代码支持,适合研究土耳其语NLP的学者与开发者使用

评估变压器、大语言模型等模型架构的性能,需要涵盖多维度的综合基准。其中,自然语言理解(NLU)评估尤为关键,是衡量模型能力的基础标准。尽管英语已有GLUE基准,中文、法语、日语也分别有CLUE、FLUE、JGLUE,但目前尚无针对土耳其语的同类基准。为此,本文提出TrGLUE——一个涵盖多种土耳其语NLU任务的综合性基准;同时发布专用于情感分析的SentiTurca基准。为支持研究,还提供了基于Transformer模型的微调与评估代码。TrGLUE采用本土化语料库,其任务形式与GLUE一致,标签通过结合强语言模型标注、跨模型一致性检查及人工验证的半自动化流程生成,确保语言自然性,减少直接翻译带来的偏差,并实现可扩展、可复现的工作流。本研究旨在建立可靠的土耳其语NLU评估框架,为研究人员提供资源,并揭示高质量半自动化数据集的构建方法。

原文摘要 · Abstract (English)

Evaluating the performance of various model architectures, such as transformers, large language models (LLMs), and other NLP systems, requires comprehensive benchmarks that measure performance across multiple dimensions. Among these, the evaluation of natural language understanding (NLU) is particularly critical as it serves as a fundamental criterion for assessing model capabilities. Thus, it is essential to establish benchmarks that enable thorough evaluation and analysis of NLU abilities from diverse perspectives. While the GLUE benchmark has set a standard for evaluating English NLU, similar benchmarks have been developed for other languages, such as CLUE for Chinese, FLUE for French, and JGLUE for Japanese. However, no comparable benchmark currently exists for the Turkish language. To address this gap, we introduce TrGLUE, a comprehensive benchmark encompassing a variety of NLU tasks for Turkish. In addition, we present SentiTurca, a specialized benchmark for sentiment analysis. To support researchers, we also provide fine-tuning and evaluation code for transformer-based models, facilitating the effective use of these benchmarks. TrGLUE comprises Turkish-native corpora curated to mirror the domains and task formulations of GLUE-style evaluations, with labels obtained through a semi-automated pipeline that combines strong LLM-based annotation, cross-model agreement checks, and subsequent human validation. This design prioritizes linguistic naturalness, minimizes direct translation artifacts, and yields a scalable, reproducible workflow. With TrGLUE, our goal is to establish a robust evaluation framework for Turkish NLU, empower researchers with valuable resources, and provide insights into generating high-quality semi-automated datasets.

语言理解情感分析土耳其语基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。