用大模型辅助构建法国维修法规知识图谱,提升法律文本结构化能力。
LLM-Assisted Ontology Engineering and Construction of a French Legal Knowledge Graph
- 分两阶段:先从文本中提取实体和关系,再基于本体构建完整知识图谱。
- 融合后实体与谓词重复率降低,90%以上类名对齐,新属性少于20%。
- 适合法律信息工程、工业维护系统等需结构化法规数据的场景。
维修法规是复杂的法律文本,难以用于具体案件处理,也难集成到实际系统中。本文提出一种基于大模型的两阶段工作流,用于构建法国维修法规的知识图谱:第一阶段基于SEMLEG核心本体,从分层语料样本中开放抽取带类型的实体与三元组,通过嵌入融合实现标签归一化,并推导候选对象属性及其签名(域与值域);第二阶段利用生成的本体引导全语料闭合抽取三元组并构建RDF图。实验使用GPT-4.1和mistral-large-2512,结果表明结构化输出稳定,类名近完全对齐,融合后实体与谓词重复显著减少。少于20%的三元组引入未知属性,而较低的精确签名符合率揭示了现有谓词的新域-值组合。这些发现表明,谓词归一化与新关系签名验证是工业维修场景下关键的优化步骤。
原文摘要 · Abstract (English)
Maintenance regulations are complex legal texts that are difficult to exploit when addressing a specific case and challenging to integrate into operational systems. This paper presents a two-stage LLM-assisted workflow for French maintenance regulations: ontology engineering from a SEMLEG-based core ontology, followed by construction of an ontology-grounded French legal knowledge graph. The first stage consists in the open extraction of typed entities and triples from a stratified corpus sample, the normalization of labels through embedding-based fusion, and the induction of candidate object properties with their signature (domain and range). The second stage uses the resulting ontology to guide the closed extraction of triples and RDF graph construction over the full corpus. Experiments with GPT-4.1 and mistral-large-2512 show robust structured outputs, near-complete class alignment, and a substantial reduction of duplicated entities and predicates after fusion. Fewer than 20% of triples introduce unseen properties, while lower exact signature compliance reveals new domain-range combinations for existing predicates. These results point to predicate normalization and the validation of newly observed relation signatures as key refinement steps for industrial maintenance settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。