arXiv:2604.10853cs.AI2026-04被引 1

构建可执行的KG评估基准,测试保险合同中知识缺口与重叠分析能力

A Benchmark for Gap and Overlap Analysis as a Test of KG Task Readiness

  • 将合同文本与形式化本体对齐,生成可审计的证据链
  • 10份保单、58个场景、带条款引用的答案标签,支持系统性对比
  • 显式建模优于纯文本大模型,提升诊断一致性和可解释性

面向任务的知识图谱(KG)质量评估正关注本体能否回答用户真正关心的问题,且具备可复现、可解释、可追溯证据的特点。本文聚焦于政策类文档(如保险合同)的缺口与重叠分析:给定一个场景,判断哪些文档支持该场景(重叠),哪些不支持(缺口),并提供可辩护的依据。此类判断源于实际覆盖范围和限制差异,而非缺失数据,因此直接检验了KG的任务就绪状态。本文提出一个可执行、可审计的基准,将自然语言合同文本与形式化本体及带证据标注的真实答案对齐,支持方法系统的比较。基准包含:(i) 经领域专家评审的10份简化但多样的人寿保险合同;(ii) 域本体(TBox)及基于合同事实实例化的知识库(ABox);(iii) 58个结构化场景,配以SPARQL查询、合同级结果和条款级引用,用于标注每项判断。通过该资源,我们对比了仅依赖文本的LLM基线与基于本体的流水线在相同场景下的表现,结果显示显式建模显著提升了一致性与诊断能力。尽管应用于缺口与重叠分析,该基准可作为通用模板,用于评估KG质量,并支撑本体学习、知识库填充与证据驱动问答等下游任务。

原文摘要 · Abstract (English)

Task-oriented evaluation of knowledge graph (KG) quality increasingly asks whether an ontology-based representation can answer the competency questions that users actually care about, in a manner that is reproducible, explainable, and traceable to evidence. This paper adopts that perspective and focuses on gap and overlap analysis for policy-like documents (e.g., insurance contracts), where given a scenario, which documents support it (overlap) and which do not (gap), with defensible justifications. The resulting gap/overlap determinations are typically driven by genuine differences in coverage and restrictions rather than missing data, making the task a direct test of KG task readiness rather than a test of missing facts or query expressiveness. We present an executable and auditable benchmark that aligns natural-language contract text with a formal ontology and evidence-linked ground truth, enabling systematic comparison of methods. The benchmark includes: (i) ten simplified yet diverse life-insurance contracts reviewed by a domain expert, (ii) a domain ontology (TBox) with an instantiated knowledge base (ABox) populated from contract facts, and (iii) 58 structured scenarios paired with SPARQL queries with contract-level outcomes and clause-level excerpts that justify each label. Using this resource, we compare a text-only LLM baseline that infers outcomes directly from contract text against an ontology-driven pipeline that answers the same scenarios over the instantiated KG, demonstrating that explicit modeling improves consistency and diagnosis for gap/overlap analyses. Although demonstrated for gap and overlap analysis, the benchmark is intended as a reusable template for evaluating KG quality and supporting downstream work such as ontology learning, KG population, and evidence-grounded question answering.

知识图谱合同分析可解释性评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。