arXiv:2601.07528cs.CLcs.AI2026-01ACL被引 10

构建伊斯兰问答基准与智能检索框架,提升大模型回答的准确性与可信度。

From RAG to Agentic RAG for Faithful Islamic Question Answering

  • 提出基于结构化工具调用的智能检索框架,实现证据迭代查找与答案修正。
  • 在3,810项双语数据上验证,小模型Qwen3 4B仍达顶尖性能并具跨语言鲁棒性。
  • 专为检测幻觉与不确定时拒绝回答设计,填补传统评估空白。

大型语言模型在伊斯兰问答中的应用日益广泛,但未经证实的回答可能带来严重宗教后果。现有基于选择题和机器阅读理解的评估无法捕捉真实场景下的关键失败模式,如自由生成式幻觉以及在证据不足时的拒答能力。为此,我们构建了IslamicFaithQA,一个包含3,810个条目的双语(阿拉伯语/英语)生成式基准,具备原子级单个正确答案,可直接衡量幻觉与拒答行为。我们还开发了一套端到端的伊斯兰知识建模体系:(i) 2.5万条阿拉伯语文本-标注的SFT推理对,(ii) 5,000条双语偏好样本用于奖励引导对齐,(iii) 约6,000个经文级别的《古兰经》检索语料库(ayat)。在此基础上,我们提出一种智能型《古兰经》接地框架(agentic RAG),通过结构化工具调用实现迭代证据搜寻与答案修订。在以阿拉伯语为中心及多语言大模型上的实验表明,检索显著提升正确率,而agentic RAG相比标准RAG带来最大提升,达到当前最优表现,并在小模型(如Qwen3 4B)上仍保持强阿拉伯语-英语鲁棒性。所有数据集已公开发布于https://huggingface.co/datasets/QCRI/IslamicFaithQA。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences. Yet standard MCQ/MRC-style evaluations (MCQ: Multiple choice questions, MRC: Machine Reading Comprehension) do not capture key real-world failure modes, notably free-form hallucinations and the ability to abstain when evidence is insufficient. To address this gap, we introduce IslamicFaithQA, a 3,810-item bilingual (Arabic/English) generative benchmark with atomic single-gold answers, which enables direct measurement of hallucination and abstention. We additionally developed an end-to-end grounded Islamic modeling suite consisting of (i) 25K Arabic text-grounded SFT reasoning pairs, (ii) 5K bilingual preference samples for reward-guided alignment, and (iii) a verse-level Qur'an retrieval corpus of ~6k atomic verses (ayat). Building on these resources, we develop an agentic Quran-grounding framework (agentic RAG) that uses structured tool calls for iterative evidence seeking and answer revision. Experiments across Arabic-centric and multilingual LLMs show that retrieval improves correctness and that agentic RAG yields the largest gains beyond standard RAG, achieving state-of-the-art performance and stronger Arabic-English robustness even with a small model (i.e., Qwen3 4B). We made the datasets are publicly available. https://huggingface.co/datasets/QCRI/IslamicFaithQA

伊斯兰AI检索增强幻觉抑制多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。