arXiv:2608.20390cs.CLcs.AI2026-08

Ansari是基于检索的伊斯兰智能助手,14万次对话验证其事实准确与价值对齐。

Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations

  • 通过检索《古兰经》、圣训等权威文献生成答案并附引用,杜绝虚构内容。
  • 在伊斯兰知识测评中超越前沿模型,法律推理任务表现优异且拒绝错误前提。
  • 适合宗教敏感场景的AI设计,为信仰类应用提供可复用的系统框架。

通用大语言模型在回答宗教问题时存在捏造经文和价值观偏差的风险。我们提出Ansari,一个已部署的检索增强型伊斯兰AI助手,自2023年6月起已处理超过14万次跨25种语言的对话。该系统采用代理式检索循环:工具型语言模型在经过认证的伊斯兰文献库(包括《古兰经》、圣训集、法理学百科全书及注释文献)中搜索,并仅基于检索结果生成答案,附带可验证引文。我们描述了其架构(代理循环、检索工具、语料库及编码编辑与神学政策的系统提示)、多平台部署(网页、移动端、WhatsApp、Model Context Protocol服务器及代理技能),以及14万次真实对话揭示的穆斯林使用模式。评估结果显示:在多项独立测试中,包括伊斯兰学术考试零样本表现、斋月期间的人工评分验证,以及两个外部基准测试,Ansari目前在公开的IslamicMMLU榜单上领先于前沿模型,在伊斯兰法律推理(IslamicLegalBench)任务中表现竞争力,且显著抑制虚假前提。研究得出普适性启示:检索是必要但非充分条件,系统提示既是技术也是神学建构,而模型形成过程中缺乏社区参与仍是核心短板。

原文摘要 · Abstract (English)

General-purpose large language models (LLMs) are increasingly used to answer religious questions, but for Islamic content they carry two serious risks: factual fabrication (inventing Qur'anic verses or hadith) and subtle value misalignment. We present Ansari, a deployed, retrieval-grounded Islamic AI assistant that has handled more than 140,000 conversations across 25+ languages since June 2023. Ansari is built around an agentic retrieval loop: a tool-using language model issues searches against authenticated Islamic corpora -- the Qur'an, hadith collections, a multi-volume jurisprudence (fiqh) encyclopedia, and exegetical (tafsir) sources -- and answers only on the basis of what it retrieves, with citations attached for verification. We describe the system's architecture (the agent loop, the retrieval tools, the corpora, and the system prompt that encodes editorial and theological policy), its multi-platform deployment (web, mobile, WhatsApp, and as a Model Context Protocol server and an Agent Skill), and what 140,000 real conversations reveal about how Muslims actually use such a tool. We report results on several complementary evaluations -- zero-shot performance on accredited institutional exams, a human-rated validation during Ramadan, and two independent, externally run benchmarks on which Ansari currently tops the public IslamicMMLU leaderboard ahead of frontier models and is competitive on Islamic legal reasoning (IslamicLegalBench) while strongly resisting false premises -- and draw out lessons that generalize beyond Islam to any faith- or values-sensitive deployment of LLMs: grounding is necessary but not sufficient, the system prompt is a theological as much as a technical artifact, and the absence of community in how models are formed remains a hard gap.

宗教AI检索增强大模型伦理伊斯兰知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。