构建首个印地语系多语言知识问答基准,评估大模型在印度本土知识上的表现。
L3Cube-IndicQuest: A Benchmark Question Answering Dataset for Evaluating Knowledge of LLMs in Indic Context
- 设计覆盖19种印地语系语言的200组问答对,涵盖5个印度特有领域。
- 提供黄金标准数据集,支持基于参考答案和大模型自评两种评估方式。
- 适合研究多语言大模型、印度本地化AI及低资源语言技术的学者使用。
大语言模型在多语言模型中已取得显著进展,融入了多种印地语系语言。然而,亟需定量评估这些语言是否与全球主导语言(如英语)表现相当。目前缺乏专门用于评估大模型在印地语系语言中区域知识掌握程度的基准数据集。本文提出L3Cube-IndicQuest,一个高质量的事实性问答基准数据集,旨在评估多语言大模型在印地语系各语言中对印度地区知识的理解能力。该数据集包含200组问答对,每组对应英文和19种印地语系语言,覆盖五个具有印度地域特色的领域。我们期望该数据集能作为基准,为评估大模型在理解与表达印度相关知识方面提供真实参考。该数据集可用于基于参考答案的评估,也可用于大模型作为裁判的自评方法。数据集已公开于https://github.com/l3cube-pune/indic-nlp。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have made significant progress in incorporating Indic languages within multilingual models. However, it is crucial to quantitatively assess whether these languages perform comparably to globally dominant ones, such as English. Currently, there is a lack of benchmark datasets specifically designed to evaluate the regional knowledge of LLMs in various Indic languages. In this paper, we present the L3Cube-IndicQuest, a gold-standard factual question-answering benchmark dataset designed to evaluate how well multilingual LLMs capture regional knowledge across various Indic languages. The dataset contains 200 question-answer pairs, each for English and 19 Indic languages, covering five domains specific to the Indic region. We aim for this dataset to serve as a benchmark, providing ground truth for evaluating the performance of LLMs in understanding and representing knowledge relevant to the Indian context. The IndicQuest can be used for both reference-based evaluation and LLM-as-a-judge evaluation. The dataset is shared publicly at https://github.com/l3cube-pune/indic-nlp .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。