arXiv:2608.20246cs.IR2026-08

针对阿拉伯语伊斯兰教法问答,构建了精准的检索测试集并验证多种策略。

What Makes a Good Fiqh Retriever? Answer Retrieval for Arabic Islamic Jurisprudence

  • 按教法学派设计检索模型,提升特定问题召回率。
  • 最佳模型在MRR@5上达0.524,微调后提升至0.553。
  • 核心挑战是区分含答案与仅话题相似的段落,适合教法研究者。

检索增强生成用于伊斯兰问答系统,但多数系统端到端评估,难以区分检索失败与生成失败。本文聚焦阿拉伯语教法(fiqh)中的答案承载检索任务,其中只有明确给出问题所求裁决的段落才算相关。我们构建了一个阿拉伯语教法检索测试集,并评估密集检索、词法检索、混合检索、微调及教派感知检索策略。最优检索器在MRR@5上得分为0.524,微调后提升至0.553。混合检索对强模型增益有限,而教派感知过滤在学派特定问题上使MRR@5翻倍。进一步误差分析显示,主要挑战在于区分含裁决的段落与仅主题相关的段落。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation is used for Islamic question answering, but most systems are evaluated end-to-end, making retrieval failures difficult to isolate from generation failures. We study answer-bearing retrieval for Arabic fiqh, where a passage is relevant only if it states the ruling required by the question. We build a retrieval test collection for Arabic fiqh and use it to evaluate dense, lexical, hybrid, fine-tuned, and madhhab-aware retrieval strategies. The best retriever achieves 0.524 MRR@5, while fine-tuning improves performance to 0.553. Hybrid retrieval provides limited gains for strong models, whereas madhhab-aware filtering more than doubles MRR@5 on school-specific questions. We further present an error analysis showing that the main challenge is distinguishing answer-bearing passages from topically similar passages that do not contain the requested ruling.

教法问答信息检索阿拉伯语NLP检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。