arXiv:2511.12387cs.CLcs.AI2025-11

首个专为泰米尔语设计的评测基准,揭示大模型在低资源语言上的理解短板。

From Phonemes to Meaning: Evaluating Large Language Models on Tamil

  • 构建首个泰米尔语专用评测集,基于斯里兰卡中小学语文考题
  • 闭源模型表现最优,开源模型普遍落后,复杂语法任务下降明显
  • 模型表现与语言类别识别能力无关,或依赖数据暴露而非真正理解

大型语言模型在高资源语言中展现出强大泛化能力,但在低资源且形态丰富的语言如泰米尔语中的语言能力仍待探索。现有多语言评测常依赖英文数据翻译,难以体现目标语言的语言文化特征。为此,我们引入ILAKKANAM——首个由人工精心构建的泰米尔语语言评测基准,包含820道来自斯里兰卡中小学泰米尔语考试的题目。每道题由受训语言学家按五个语言类别和一个事实知识类别标注,覆盖小学一年级至十三年级,确保广泛的语言覆盖。我们采用标准化框架评估闭源与开源模型。结果显示,Gemini 2.5整体表现最佳,而开源模型明显落后,凸显语言基础差距。分类别、分年级分析表明,所有模型在低年级题目上表现良好,但随着语言复杂度提升,性能显著下降。进一步发现,模型整体表现与其识别语言类别的能力无强相关性,暗示其表现可能源于数据曝光而非真实理解。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown strong generalization across tasks in high-resource languages; however, their linguistic competence in low-resource and morphologically rich languages such as Tamil remains largely unexplored. Existing multilingual benchmarks often rely on translated English datasets, failing to capture the linguistic and cultural nuances of the target language. To address this gap, we introduce ILAKKANAM, the first Tamil-specific linguistic evaluation benchmark manually curated using 820 questions from Sri Lankan school-level Tamil subject examination papers. Each question is annotated by trained linguists under five linguistic categories and a factual knowledge category, spanning Grades 1--13 to ensure broad linguistic coverage. We evaluate both closed-source and open-source LLMs using a standardized evaluation framework. Our results show that Gemini 2.5 achieves the highest overall performance, while open-source models lag behind, highlighting the gap in linguistic grounding. Category- and grade-wise analyses reveal that all models perform well on lower-grade questions but show a clear decline as linguistic complexity increases. Further, no strong correlation is observed between a model's overall performance and its ability to identify linguistic categories, suggesting that performance may be driven by exposure rather than genuine understanding.

泰米尔语大模型评测低资源语言语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。