评估大模型在宗教问题上的可靠性与拒答能力,发现英文表现优于阿拉伯文。
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
- 构建针对四大逊尼派教法学派的FiqhQA基准,支持中英双语
- GPT-4o准确率最高,Gemini和Fanar拒答行为更优
- 所有模型在阿拉伯语上表现下降,凸显多语言宗教推理短板
尽管大语言模型(LLMs)在多个领域被广泛使用,其在宗教领域的可靠性与准确性仍缺乏系统评估。本文提出一个新基准FiqhQA,专门用于评估大模型生成伊斯兰教法裁决的能力,涵盖四大逊尼派教法学派,支持阿拉伯语和英语。与以往研究不同,本工作不仅关注准确率,还重点评估模型在无法回答时的拒答能力。零样本与拒答实验显示,不同模型、语言及教法学派间存在显著差异:GPT-4o在准确率上领先,而Gemini与Fanar在拒答行为上表现更佳,有助于减少自信错误答案。值得注意的是,所有模型在阿拉伯语上的性能均明显下降,反映出非英语环境下宗教推理能力的局限性。据我们所知,这是首个针对细粒度伊斯兰教法学派裁决生成并评估拒答行为的研究。结果强调了在宗教应用中进行任务特定评估与审慎部署的重要性。
原文摘要 · Abstract (English)
Despite the increasing usage of Large Language Models (LLMs) in answering questions in a variety of domains, their reliability and accuracy remain unexamined for a plethora of domains including the religious domains. In this paper, we introduce a novel benchmark FiqhQA focused on the LLM generated Islamic rulings explicitly categorized by the four major Sunni schools of thought, in both Arabic and English. Unlike prior work, which either overlooks the distinctions between religious school of thought or fails to evaluate abstention behavior, we assess LLMs not only on their accuracy but also on their ability to recognize when not to answer. Our zero-shot and abstention experiments reveal significant variation across LLMs, languages, and legal schools of thought. While GPT-4o outperforms all other models in accuracy, Gemini and Fanar demonstrate superior abstention behavior critical for minimizing confident incorrect answers. Notably, all models exhibit a performance drop in Arabic, highlighting the limitations in religious reasoning for languages other than English. To the best of our knowledge, this is the first study to benchmark the efficacy of LLMs for fine-grained Islamic school of thought specific ruling generation and to evaluate abstention for Islamic jurisprudence queries. Our findings underscore the need for task-specific evaluation and cautious deployment of LLMs in religious applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。