用历史词典增强阿拉伯语大模型,提升古籍理解准确率
Grounding Arabic LLMs in the Doha Historical Dictionary: Retrieval-Augmented Understanding of Quran and Hadith
- 基于多时态词典检索,精准获取宗教文本的历史语义
- 使本土模型准确率达85%以上,接近商用模型表现
- 适合研究伊斯兰文献或阿拉伯语历史语言的学者
大语言模型在诸多语言任务中取得显著进展,但在处理《古兰经》和圣训等复杂历史宗教阿拉伯文文本时仍存在困难。为此,本文构建了一个基于时间演变词典知识的检索增强生成(RAG)框架。不同于依赖通用语料库的现有系统,本方法从《多哈阿拉伯语历史词典》(DHDA)中检索证据,该资源记录了阿拉伯语词汇的历史演变。所提流程结合混合检索与意图导向路由机制,为大模型提供精确、上下文相关的史学信息。实验表明,该方法使本地阿拉伯语大模型(如Fanar和ALLaM)的准确率超过85%,显著缩小与商用模型Gemini的差距。Gemini同时作为模型裁判用于自动评估,自动化判断经人工验证,一致性达0.87(kappa)。错误分析揭示了符号标记与复合表达等关键语言挑战。结果证明,在RAG框架中整合历时词典资源能有效提升阿拉伯语历史宗教文本的理解能力。代码与资源已公开:https://github.com/somayaeltanbouly/Doha-Dictionary-RAG。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable progress in many language tasks, yet they continue to struggle with complex historical and religious Arabic texts such as the Quran and Hadith. To address this limitation, we develop a retrieval-augmented generation (RAG) framework grounded in diachronic lexicographic knowledge. Unlike prior RAG systems that rely on general-purpose corpora, our approach retrieves evidence from the Doha Historical Dictionary of Arabic (DHDA), a large-scale resource documenting the historical development of Arabic vocabulary. The proposed pipeline combines hybrid retrieval with an intent-based routing mechanism to provide LLMs with precise, contextually relevant historical information. Our experiments show that this approach improves the accuracy of Arabic-native LLMs, including Fanar and ALLaM, to over 85\%, substantially reducing the performance gap with Gemini, a proprietary large-scale model. Gemini also serves as an LLM-as-a-judge system for automatic evaluation in our experiments. The automated judgments were verified through human evaluation, demonstrating high agreement (kappa = 0.87). An error analysis further highlights key linguistic challenges, including diacritics and compound expressions. These findings demonstrate the value of integrating diachronic lexicographic resources into retrieval-augmented generation frameworks to enhance Arabic language understanding, particularly for historical and religious texts. The code and resources are publicly available at: https://github.com/somayaeltanbouly/Doha-Dictionary-RAG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。