构建首个超大规模伊斯兰经文问答数据集,助力精准理解宗教文本。
A Benchmark Dataset with Larger Context for Non-Factoid Question Answering over Islamic Text
- 构建7.3万条带上下文的经文问答对,覆盖古兰经注释与圣训
- 发现现有自动评估指标与专家判断差异巨大,模型一致性仅11%-20%
- 适合研究宗教文本理解、多轮问答与评估方法的学者使用
在数字时代,高效准确地访问和理解宗教文本(尤其是《古兰经》及圣训)亟需先进的问答系统。然而,针对古兰经注释与圣训中细节性问题的问答系统仍严重匮乏。为此,本文构建了一个面向古兰经注释与圣训领域的综合性问答数据集,包含超过73,000个问答对,是该领域迄今最大规模的数据集。所有问题与答案均附有详细上下文信息,可作为训练与评估专用问答系统的宝贵资源。尽管本研究建立了该领域的基准评测体系,但后续的人工评估揭示了现有自动评价方法的重大局限:模型判断与专家一致率仅为11%至20%,而上下文理解能力在50%至90%之间波动。这一差距凸显了需要更精细的评估方法来捕捉宗教文本理解的复杂性,超越传统自动指标的不足。
原文摘要 · Abstract (English)
Accessing and comprehending religious texts, particularly the Quran (the sacred scripture of Islam) and Ahadith (the corpus of the sayings or traditions of the Prophet Muhammad), in today's digital era necessitates efficient and accurate Question-Answering (QA) systems. Yet, the scarcity of QA systems tailored specifically to the detailed nature of inquiries about the Quranic Tafsir (explanation, interpretation, context of Quran for clarity) and Ahadith poses significant challenges. To address this gap, we introduce a comprehensive dataset meticulously crafted for QA purposes within the domain of Quranic Tafsir and Ahadith. This dataset comprises a robust collection of over 73,000 question-answer pairs, standing as the largest reported dataset in this specialized domain. Importantly, both questions and answers within the dataset are meticulously enriched with contextual information, serving as invaluable resources for training and evaluating tailored QA systems. However, while this paper highlights the dataset's contributions and establishes a benchmark for evaluating QA performance in the Quran and Ahadith domains, our subsequent human evaluation uncovered critical insights regarding the limitations of existing automatic evaluation techniques. The discrepancy between automatic evaluation metrics, such as ROUGE scores, and human assessments became apparent. The human evaluation indicated significant disparities: the model's verdict consistency with expert scholars ranged between 11% to 20%, while its contextual understanding spanned a broader spectrum of 50% to 90%. These findings underscore the necessity for evaluation techniques that capture the nuances and complexities inherent in understanding religious texts, surpassing the limitations of traditional automatic metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。