用多阶段训练打造双语伊斯兰检索模型,提升跨语言信息查找效果
Multi-stage Training of Bilingual Islamic LLM for Neural Passage Retrieval
- 基于XLM-R做轻量化双语模型,分阶段训练融合通用与伊斯兰领域数据
- 在MS MARCO和自建英文伊斯兰语料上训练,检索性能超越单语模型
- 适合需要跨语言伊斯兰文本检索的研究者与应用开发者
本研究聚焦伊斯兰领域自然语言处理,致力于构建一个伊斯兰神经检索模型。通过利用强大的XLM-R模型,并采用语言缩减技术,开发出轻量级双语大语言模型。针对伊斯兰领域特有的挑战——域内语料主要为阿拉伯语,其他语言(包括英语)资源有限,提出一种多阶段训练方法,结合大规模通用检索数据集(如MS MARCO)与小规模域内数据集,以提升检索性能。此外,通过数据增强技术并依托可靠伊斯兰来源,构建了英文域内检索数据集,进一步丰富了领域专用数据。实验表明,结合领域适配与多阶段训练的双语伊斯兰神经检索模型,在下游检索任务中优于单语模型。
原文摘要 · Abstract (English)
This study examines the use of Natural Language Processing (NLP) technology within the Islamic domain, focusing on developing an Islamic neural retrieval model. By leveraging the robust XLM-R model, the research employs a language reduction technique to create a lightweight bilingual large language model (LLM). Our approach for domain adaptation addresses the unique challenges faced in the Islamic domain, where substantial in-domain corpora exist only in Arabic while limited in other languages, including English. The work utilizes a multi-stage training process for retrieval models, incorporating large retrieval datasets, such as MS MARCO, and smaller, in-domain datasets to improve retrieval performance. Additionally, we have curated an in-domain retrieval dataset in English by employing data augmentation techniques and involving a reliable Islamic source. This approach enhances the domain-specific dataset for retrieval, leading to further performance gains. The findings suggest that combining domain adaptation and a multi-stage training method for the bilingual Islamic neural retrieval model enables it to outperform monolingual models on downstream retrieval tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。