arXiv:2411.02538cs.CL2024-11被引 27

首个面向印度语系语言的多任务理解评测基准,填补低资源语言评估空白。

MILU: A Multi-task Indic Language Understanding Benchmark

  • 构建覆盖11种印度语、41个主题的多任务评测集,涵盖文化与通用知识。
  • 主流模型平均准确率仅74%,文化相关领域表现显著低于科技类。
  • 开源所有数据和代码,助力多语言模型公平评估与研究。

评估大语言模型在低资源及语言多样性语言中的能力仍是自然语言处理的重要挑战,尤其对使用非拉丁字母的语言(如印度地区语言)而言。现有评测基准主要集中于英语,难以全面评估模型在这些语言中的表现。本文提出MILU——一个面向印度语系语言的多任务理解评测基准,涵盖11种印度语言、8个领域、41个主题,内容兼具通用知识与地方文化特色,包括地区考试题源、本地历史、艺术、节日与法律等。我们评估了42个大语言模型,发现当前模型在该基准上整体表现不佳,其中GPT-4o表现最佳,平均准确率为74%。开放多语言模型优于语言特化微调模型,后者仅略高于随机猜测水平。模型在高资源语言中表现更好,且在人文、法律等文化相关领域表现显著落后于科学与数学等通用领域。据我们所知,MILU是首个专注于印度语系语言的综合性评测基准,为实现更全面的文化适配性评估迈出关键一步。所有代码、数据与工具均已公开,推动开放研究。

原文摘要 · Abstract (English)

Evaluating Large Language Models (LLMs) in low-resource and linguistically diverse languages remains a significant challenge in NLP, particularly for languages using non-Latin scripts like those spoken in India. Existing benchmarks predominantly focus on English, leaving substantial gaps in assessing LLM capabilities in these languages. We introduce MILU, a Multi task Indic Language Understanding Benchmark, a comprehensive evaluation benchmark designed to address this gap. MILU spans 8 domains and 41 subjects across 11 Indic languages, reflecting both general and culturally specific knowledge. With an India-centric design, incorporates material from regional and state-level examinations, covering topics such as local history, arts, festivals, and laws, alongside standard subjects like science and mathematics. We evaluate over 42 LLMs, and find that current LLMs struggle with MILU, with GPT-4o achieving the highest average accuracy at 74 percent. Open multilingual models outperform language-specific fine-tuned models, which perform only slightly better than random baselines. Models also perform better in high resource languages as compared to low resource ones. Domain-wise analysis indicates that models perform poorly in culturally relevant areas like Arts and Humanities, Law and Governance compared to general fields like STEM. To the best of our knowledge, MILU is the first of its kind benchmark focused on Indic languages, serving as a crucial step towards comprehensive cultural evaluation. All code, benchmarks, and artifacts are publicly available to foster open research.

多语言评估印度语系文化理解大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。