arXiv:2504.09714cs.CLcs.AI2025-04被引 14

评估17个土耳其语数据集质量,发现七成不达标,强调低资源语言数据需更严格筛选。

Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish

  • 从6个维度评估17个土耳其语数据集,结合人工与大模型评分。
  • 70%数据集未达质量标准,85%的技术术语使用不正确。
  • 大模型虽可用,但人类在文化常识判断上仍更可靠。

依赖英语或多语言资源翻译或改编的数据集,常面临语言与文化适配问题。本研究针对这一痛点,评估了17个常用土耳其语基准数据集的质量。采用涵盖六项标准的综合框架,由人工与大模型标注者进行详细评估,识别数据集优劣。结果显示,70%的基准数据集未能满足预设质量标准;技术术语使用正确性为最强指标,但仍有85%的条目不达标。尽管大模型具备潜力,其表现仍逊于人类,尤其在理解文化常识和流畅文本解读方面。GPT-4o在语法与技术任务上表现更佳,而Llama3.3-70B在正确性与文化知识评估中更突出。研究强调,低资源语言数据集的构建与适配亟需更严格的质控机制。

原文摘要 · Abstract (English)

The reliance on translated or adapted datasets from English or multilingual resources introduces challenges regarding linguistic and cultural suitability. This study addresses the need for robust and culturally appropriate benchmarks by evaluating the quality of 17 commonly used Turkish benchmark datasets. Using a comprehensive framework that assesses six criteria, both human and LLM-judge annotators provide detailed evaluations to identify dataset strengths and shortcomings. Our results reveal that 70% of the benchmark datasets fail to meet our heuristic quality standards. The correctness of the usage of technical terms is the strongest criterion, but 85% of the criteria are not satisfied in the examined datasets. Although LLM judges demonstrate potential, they are less effective than human annotators, particularly in understanding cultural common sense knowledge and interpreting fluent, unambiguous text. GPT-4o has stronger labeling capabilities for grammatical and technical tasks, while Llama3.3-70B excels at correctness and cultural knowledge evaluation. Our findings emphasize the urgent need for more rigorous quality control in creating and adapting datasets for low-resource languages.

数据集评估低资源语言文化适配大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。