大模型重塑圣训计算,但方法仍需专家验证与知识融合。
Hadith computational science in the age of large language models: a critical narrative review
- 用大模型和检索增强框架提升圣训数据处理能力
- 实现多语言访问与大规模语料扩展,但跨数据集比较困难
- 强调圣训真实性与权威性需专家参与,非单纯模型性能
本文审视变压器模型、检索增强流程与大语言模型(LLMs)如何重塑圣训计算科学。现有综述虽记录文献增长,却未批判性评估哪些进展方法稳健、哪些仍受基准限制、哪些未解难题阻碍学术应用。本文通过批判性叙事综述,结合对已有综述的批评、代表性原始研究的逐篇评估,以及伊斯兰学者与领域专家对真实性、权威性与负责任使用的观点整合,发现进展不均:数据资源扩展,分段任务成熟,说者与来源验证问题更形式化,且现有多语言支持、全库扩展与基于证据的评估已可通过大模型辅助实现。然而,受限于狭隘语料、弱可比基准、合成到真实迁移差距、说者身份识别困难、预处理脆弱性、可复现性差及专家验证稀缺。重要缺口存在于主流基准之外:非正统与罕见语料、注释与解释文献、与古兰经及先知传记的跨源关联,以及教法学支持证据。我们主张将圣训计算视为证据基础设施问题,需知识整合、溯源追踪与专家监督。据此提出研究议程,以强化该领域的方法论并提升对伊斯兰学术的支持。
原文摘要 · Abstract (English)
We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, and large language models (LLMs). Recent reviews document growth in the literature, but they do not yet provide a critical account of which advances are methodologically robust, which remain benchmark-bound, and which unresolved problems still limit scholarly use. We address this gap through a critical narrative review that combines critique of existing reviews, paper-level appraisal of representative original studies, and synthesis of Islamic scholar and domain-expert perspectives on authenticity, authority, and responsible use. We find uneven progress. Data resources have expanded, segmentation tasks have matured, narrator and source-verification problems are better formalized, and LLM-assisted workflows now support corpus-scale enrichment, multilingual access, and grounded evaluation. At the same time, progress remains constrained by narrow corpora, weak benchmark comparability, synthetic-to-real transfer gaps, narrator identity resolution, preprocessing fragility, limited reproducibility, and sparse expert-grounded validation. We show that important gaps lie beyond dominant benchmarks: non-canonical and obscure corpora, commentary and explanatory literature, cross-source links with Qur'an and seerah, and fiqh-facing evidence support. We argue that hadith computation should be assessed less as isolated model performance than as an evidence infrastructure problem requiring knowledge integration, provenance, and expert supervision. On this basis, we define a research agenda for making the field methodologically stronger and more useful to Islamic scholarship.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。