arXiv:2608.07530cs.AIcs.CL2026-08中稿 · ISWC 2026

构建首个自然语言转SHACL的评测基准,助力知识图谱验证自动化。

NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation

论文配图:NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation
图 1 · 摘自论文原文
  • 设计多维度评测集,覆盖复杂逻辑与结构模式的自然语言到SHACL转换。
  • 评估4个主流大模型,发现其生成语法正确但语义匹配度不足。
  • 适合知识图谱、语义网和AI可解释性研究者使用。

SHACL是验证RDF知识图谱一致性的重要技术,但编写SHACL需要专业知识,多数领域专家难以胜任。将自然语言需求转化为SHACL(NL2SHACL)可降低门槛。然而,目前缺乏专用评测基准,且生成结果需超越字符串比对的语义评估方法,因语义等价的规则可能在序列化和结构上不同。为此,我们提出NL2SHACL-Bench,首个面向自然语言到SHACL转换的基准套件。利用该套件,我们评估了四个前沿大语言模型在此任务上的表现。结果表明,当前模型能生成语法正确的SHACL,但在复杂逻辑与结构模式上仍难以生成语义等价的约束。这说明NL2SHACL-Bench为衡量该领域进展提供了有效基础。

原文摘要 · Abstract (English)

SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical expertise that most domain experts lack. Translating natural language requirements into SHACL (NL2SHACL) would lower this barrier. However, there is no dedicated benchmark for NL2SHACL, and evaluating generated shapes requires methods beyond string comparison, as semantically equivalent shapes can differ in serialisation and structure. To tackle these challenges, we present NL2SHACL-Bench, a benchmark suite for natural language to SHACL translation. Using NL2SHACL-Bench, we evaluate four state-of-the-art large language models (LLMs) for this task. Our results show that current LLMs are highly capable of generating syntactically valid SHACL, but still struggle to produce semantically equivalent constraints for complex logical and structural patterns. This indicates that NL2SHACL-Bench provides a meaningful basis for measuring advances in the NL2SHACL state of the art.

知识图谱自然语言SHACL评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。