为13种印度语言构建了首个大规模RAG评估与训练数据集
IndicRAGSuite: Large-Scale Datasets and a Benchmark for Indian Language RAG Systems
- 手动翻译MS MARCO数据集,创建多语言检索与生成基准
- 利用LLM从19种印度语维基构建超百万级问答三元组数据
- 适合研究多语言AI、跨语言信息检索及印度语NLP的开发者
检索增强生成(RAG)系统使语言模型能访问相关信息,生成准确且有依据的回答。然而,由于缺乏两个关键资源——(1)检索与生成任务的评估基准,(2)多语言检索的大规模训练数据,印度语言的高质量RAG系统发展受限。现有基准和数据集多集中于英语或高资源语言,难以覆盖印度多样化的语言生态。为此,我们构建了IndicMSMarco,一个涵盖13种印度语言的多语言基准,通过人工翻译MS MARCO-dev集中的1000个多样化查询生成。同时,我们利用先进大模型,从19种印度语言的维基百科中构建大规模(问题, 答案, 相关段落)三元组数据集。此外,还包含原始MS MARCO数据集的翻译版本,以丰富训练数据并确保与真实信息检索任务对齐。相关资源已公开:https://huggingface.co/collections/ai4bharat/indicragsuite-683e7273cb2337208c8c0fcb
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) systems enable language models to access relevant information and generate accurate, well-grounded, and contextually informed responses. However, for Indian languages, the development of high-quality RAG systems is hindered by the lack of two critical resources: (1) evaluation benchmarks for retrieval and generation tasks, and (2) large-scale training datasets for multilingual retrieval. Most existing benchmarks and datasets are centered around English or high-resource languages, making it difficult to extend RAG capabilities to the diverse linguistic landscape of India. To address the lack of evaluation benchmarks, we create IndicMSMarco, a multilingual benchmark for evaluating retrieval quality and response generation in 13 Indian languages, created via manual translation of 1000 diverse queries from MS MARCO-dev set. To address the need for training data, we build a large-scale dataset of (question, answer, relevant passage) tuples derived from the Wikipedias of 19 Indian languages using state-of-the-art LLMs. Additionally, we include translated versions of the original MS MARCO dataset to further enrich the training data and ensure alignment with real-world information-seeking tasks. Resources are available here: https://huggingface.co/collections/ai4bharat/indicragsuite-683e7273cb2337208c8c0fcb
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。