arXiv:2508.13526cs.CL2025-08中稿 · LREC 2026

评测大模型在泰卢固语中的语言能力,发现其依赖答题位置等表面线索。

MATA: Mindful Assessment of the Telugu Abilities of Large Language Models

  • 构建729题泰卢固语多选与开放题评估集,覆盖多样语言维度。
  • 11个模型表现参差,多选题依赖答案位置和干扰项模式。
  • 对比人工评分与模型自评,验证其在低资源语言中可靠性有限。

本文提出MATA,一个用于评估大语言模型在泰卢固语中能力的新评测数据集,包含729道精心设计的多选题和开放题,覆盖多种语言层面。我们在该数据集上评估了11个开源与闭源大模型,并进行了细粒度性能分析。结果表明,模型在多选题中倾向于依赖表面启发式策略,如答案位置和干扰项分布。此外,我们还比较了大模型作为裁判与人类评估在开放题上的表现,验证其在低资源语言中的可靠性。我们认为这种细粒度评估对理解模型局限性至关重要,可推动更具备语言能力的大模型发展,并为泰卢固语自然语言处理研究奠定基础。数据集已公开于:https://huggingface.co/datasets/TeluguLLMResearch/MATA。

原文摘要 · Abstract (English)

In this paper, we introduce MATA, a novel evaluation dataset to assess the ability of Large Language Models (LLMs) in Telugu language, comprising 729 carefully curated multiple-choice and open-ended questions that span diverse linguistic dimensions. We evaluate 11 open-weight and closed-source LLMs on our dataset and present a fine-grained analysis of their performance. Further, we empirically show how LLMs rely on superficial heuristics such as answer position and distractor patterns for multiple-choice questions. Finally, we also compare LLM-as-a-judge evaluation with human evaluation for open-ended questions assess its reliability in a low-resource language. We argue that such fine-grained evaluation is essential for understanding model limitations and can inform the development of more linguistically capable LLMs, while also serving as a foundation for future research in Telugu NLP. Our dataset is available at: https://huggingface.co/datasets/TeluguLLMResearch/MATA

泰卢固语大模型评测多选题分析低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。