扩充QA数据集并微调模型,提升古阿拉伯语问答系统准确率
Optimized Quran Passage Retrieval Using an Expanded QA Dataset and Fine-Tuned Language Models
- 扩展原始数据至629个问题,分三类答案类型
- AraBERT-base模型达MAP@10=0.36,MRR=0.59,提升超60%
- 对无答案问题处理成功率从25%升至75%,适合宗教文本研究者
理解《古兰经》深层含义并弥合现代标准阿拉伯语与古典阿拉伯语之间的语言鸿沟,是提升《古兰经》问答系统性能的关键。原2023年《古兰经》问答共享任务数据集仅有251个问题,且模型检索能力弱。本文重新审阅并扩充该数据集至629个问题,通过问题多样化与重述实现扩展,形成包含1895个样本的综合性数据集,涵盖单答案、多答案和零答案三类。实验对比了AraBERT、RoBERTa、CAMeLBERT、AraELECTRA及BERT等多种Transformer模型。最优模型AraBERT-base在测试中取得MAP@10=0.36、MRR=0.59,相比基线(MAP@10: 0.22,MRR: 0.37)分别提升63%和59%。此外,新方法在处理“无答案”情形时成功率达75%,高于基线的25%。结果表明,数据集优化与模型架构改进能显著提升《古兰经》问答系统的准确率、召回率与精确率。
原文摘要 · Abstract (English)
Understanding the deep meanings of the Qur'an and bridging the language gap between modern standard Arabic and classical Arabic is essential to improve the question-and-answer system for the Holy Qur'an. The Qur'an QA 2023 shared task dataset had a limited number of questions with weak model retrieval. To address this challenge, this work updated the original dataset and improved the model accuracy. The original dataset, which contains 251 questions, was reviewed and expanded to 629 questions with question diversification and reformulation, leading to a comprehensive set of 1895 categorized into single-answer, multi-answer, and zero-answer types. Extensive experiments fine-tuned transformer models, including AraBERT, RoBERTa, CAMeLBERT, AraELECTRA, and BERT. The best model, AraBERT-base, achieved a MAP@10 of 0.36 and MRR of 0.59, representing improvements of 63% and 59%, respectively, compared to the baseline scores (MAP@10: 0.22, MRR: 0.37). Additionally, the dataset expansion led to improvements in handling "no answer" cases, with the proposed approach achieving a 75% success rate for such instances, compared to the baseline's 25%. These results demonstrate the effect of dataset improvement and model architecture optimization in increasing the performance of QA systems for the Holy Qur'an, with higher accuracy, recall, and precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。