用伊斯兰圣训学方法给AI知识传递链打分,提升可信度判断。
Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems
- 借鉴圣训学的传述链与人物信誉评估体系,构建可评分的知识传递框架。
- 在2万条物理课本命题上验证,弱链接隔离和独立链互证有效。
- 适合关注AI知识可信度、可解释性的研究者与系统设计者。
现代多智能体知识系统越来越多地通过自主转换链积累知识,而非直接检索。现有溯源工作记录事件过程(执行轨迹、工具调用、证据链接),来源可靠性评估也已有成熟方法(真相发现、声誉系统)。但缺乏一个能为每条主张的传输链赋予分级、领域特定的传述人可信度的实用框架,该框架需具备完整性语义、类型化变换聚合、内容批评与服务/审查/隔离路由分离等特性。古典伊斯兰圣训科学曾面临类似问题:如何判断通过人类传述链传递的知识是否可信。历经数百年发展出严谨方法——传述链(isnad,每个主张附带完整传输链)、人物评级(rijal,系统性评估每位传述人诚信与精确度)、最弱环节评估、独立链交叉验证,以及对内容本身(matn)的独立批判。本文将这一方法论迁移至AI系统设计。我们贡献了:将圣训科学概念形式化映射到多智能体流水线的方法;实现主张链与分级传述人注册表的关系模式;结合链质量与内容批判的决策矩阵;在2万条真实物理教科书命题上的评估。结果验证了最弱链隔离与独立链交叉验证的有效性;报告了评级恢复环的部分失败(未能识别最高故障传述人);两项分析结果不明确,包括框架无法与基准内容批评达到匹配覆盖率的比较。论文始终明确指出哪些主张有证据支持,哪些尚未支持。
原文摘要 · Abstract (English)
Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (truth discovery, reputation systems). What is missing is an operational framework that attaches graded, per-domain transmitter reliability to claim-level transmission chains, with completeness semantics, transformation-typed aggregation, decoupled content criticism, and serve/review/quarantine routing. Classical Islamic hadith science confronted a structurally similar problem: deciding whether knowledge transmitted through chains of human narrators should be accepted. Over centuries it developed a rigorous methodology - isnad (a complete transmission chain attached to every claim), rijal (systematic grading of each narrator's integrity and precision), weakest-link chain evaluation, corroboration through independent chains, and matn criticism (content evaluated independently of chain quality). This paper transfers that methodology to AI system design. We contribute a formal mapping from hadith-science concepts to multi-agent pipelines, a relational schema implementing claim chains and a graded narrator registry, a decision matrix combining chain grade with content criticism, and an evaluation on 20,000 claims from real physics textbooks. The evaluation validates weakest-link quarantine and independent-chain corroboration; reports a partial failure of the grade-recovery loop, which missed the highest-fault narrator; and reports two analyses as inconclusive, including a matched-coverage comparison the framework could not reach with the reference content critic. The paper is explicit throughout about which claims the evidence does and does not yet support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。