arXiv:2510.24438cs.CLcs.AI2025-10中稿 · NeurIPS被引 4

用双代理框架评估大模型生成伊斯兰内容的准确性与一致性。

Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content

  • 设计双代理评估系统,量化验证引文与定性对比内容质量。
  • GPT-4o在宗教准确性和引文方面得分最高,达3.93和3.38。
  • 研究呼吁建立以穆斯林视角为中心的社区基准,适用于高风险领域。

大型语言模型越来越多地用于提供伊斯兰指导,但存在误引文本、误用教法或产生文化不一致回应的风险。我们对GPT-4o、Ansari AI和Fanar在真实伊斯兰博客提问下的表现进行了试点评估。采用双代理框架:定量代理负责引文验证及六维评分(如结构、伊斯兰一致性、引文),定性代理进行五维并列比较(如语气、深度、原创性)。GPT-4o在伊斯兰准确性(3.93)和引文(3.38)上得分最高,Ansari AI次之(3.68, 3.32),Fanar表现较弱(2.76, 1.82)。尽管整体表现良好,模型仍难以可靠生成准确的伊斯兰内容与引文——这在信仰敏感写作中至关重要。GPT-4o平均定量得分最高(3.90/5),Ansari AI在定性两两对比中胜出116次(共200次)。Fanar虽落后,但在伊斯兰与阿拉伯语语境下引入创新。本研究强调需建立以穆斯林视角为核心的社区驱动基准,为伊斯兰知识及其他高风险领域(如医学、法律、新闻)的可靠AI发展迈出早期一步。

原文摘要 · Abstract (English)

Large language models are increasingly used for Islamic guidance, but risk misquoting texts, misapplying jurisprudence, or producing culturally inconsistent responses. We pilot an evaluation of GPT-4o, Ansari AI, and Fanar on prompts from authentic Islamic blogs. Our dual-agent framework uses a quantitative agent for citation verification and six-dimensional scoring (e.g., Structure, Islamic Consistency, Citations) and a qualitative agent for five-dimensional side-by-side comparison (e.g., Tone, Depth, Originality). GPT-4o scored highest in Islamic Accuracy (3.93) and Citation (3.38), Ansari AI followed (3.68, 3.32), and Fanar lagged (2.76, 1.82). Despite relatively strong performance, models still fall short in reliably producing accurate Islamic content and citations -- a paramount requirement in faith-sensitive writing. GPT-4o had the highest mean quantitative score (3.90/5), while Ansari AI led qualitative pairwise wins (116/200). Fanar, though trailing, introduces innovations for Islamic and Arabic contexts. This study underscores the need for community-driven benchmarks centering Muslim perspectives, offering an early step toward more reliable AI in Islamic knowledge and other high-stakes domains such as medicine, law, and journalism.

大模型评估伊斯兰AI引文验证双代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。