arXiv:2509.00140cs.SEcs.AI2025-09中稿 · , 2026

用大模型从软件标准中自动提取术语关系,构建知识图谱。

LLM-based Zero-shot Triple Extraction for Automated Ontology Generation from Software Engineering Standards

  • 借助大模型进行零样本关系抽取,结合分段与术语归一化流程。
  • 在三个粒度上评估,效果接近且可能优于传统开放信息抽取方法。
  • 适合需要自动化构建领域知识库的开发与研究团队。

本研究提出一种基于大语言模型(LLM)的零样本关系三元组抽取方法,用于从软件工程标准(SES)中自动构建知识图谱。这些标准包含大量非结构化文本和领域术语,噪声高。本文设计了一个包含文档分段、候选术语挖掘、基于LLM的关系推理、术语归一化及跨章节对齐的完整工作流。构建了三个粒度的专家标注参考集用于评估生成的本体。实验结果表明,该方法在三元组抽取上的表现可媲美甚至优于现有的OpenIE方法,展现出在自动化本体生成中的潜力。

原文摘要 · Abstract (English)

Ontologies have supported knowledge representation and white-box reasoning for decades; thus, the automated ontology generation (AOG) plays a crucial role in scaling their use. Software engineering standards (SES) consist of long, unstructured text (with high noise) and paragraphs with domain-specific terms. In this setting, relation triple extraction (RTE), together with term extraction, constitutes the first stage toward AOG. This work proposes an open-source large language model (LLM)-assisted approach to RTE for SES. Instead of solely relying on prompt-engineering-based methods, this study promotes the use of LLMs as an aid in constructing ontologies and explores an effective AOG workflow that includes document segmentation, candidate term mining, LLM-based relation inference, term normalization, and cross-section alignment. Expert-annotated reference sets at three granularities are constructed and used to evaluate the ontology generated from the study. The results show that it is comparable and potentially superior to the OpenIE method of triple extraction.

知识图谱大模型应用自动抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。