arXiv:2507.08432cs.DBcs.CL2025-07中稿 · publication in the…被引 3

用大模型让数据校验报告变得易懂,支持多语言解释违规原因。

xpSHACL: Explainable SHACL Validation using Retrieval-Augmented Generation and Large Language Models

  • 结合规则树与大模型生成可读解释,提升非技术用户理解力。
  • 通过违规知识图谱缓存解释,提升效率与一致性。
  • 适合需解释数据验证结果的开发者、业务人员和数据治理者。

形状约束语言(SHACL)是验证RDF数据的强大工具。随着知识图谱在工业界日益受到关注,更多用户需要正确验证链接数据。然而,传统SHACL验证引擎通常以简略的英文报告形式输出,非技术人员难以理解并采取行动。本文提出xpSHACL,一个可解释的SHACL验证系统,通过将基于规则的证明树与检索增强生成(RAG)及大语言模型(LLMs)结合,生成详细、多语言、人类可读的约束违规解释。xpSHACL的关键特性在于使用违规知识图谱(Violation KG)缓存并重用解释,从而提高效率和一致性。

原文摘要 · Abstract (English)

Shapes Constraint Language (SHACL) is a powerful language for validating RDF data. Given the recent industry attention to Knowledge Graphs (KGs), more users need to validate linked data properly. However, traditional SHACL validation engines often provide terse reports in English that are difficult for non-technical users to interpret and act upon. This paper presents xpSHACL, an explainable SHACL validation system that addresses this issue by combining rule-based justification trees with retrieval-augmented generation (RAG) and large language models (LLMs) to produce detailed, multilanguage, human-readable explanations for constraint violations. A key feature of xpSHACL is its usage of a Violation KG to cache and reuse explanations, improving efficiency and consistency.

知识图谱可解释性大模型数据验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。