arXiv:2602.00699cs.AIcs.CL2026-02被引 2

用大模型从文本自动提取铸造领域术语与关系,省去人工标注。

From Prompt to Graph: Comparing LLM-Based Information Extraction Strategies in Domain-Specific Ontology Development

  • 对比三种大模型方法:预训练、上下文学习、微调,选最优策略。
  • 在有限数据下构建铸造领域本体,专家验证后准确率显著提升。
  • 适合需要快速构建专业本体的科研或工业团队使用。

本体是结构化领域知识的关键,有助于知识的可访问性、共享与复用。然而,传统本体构建依赖人工标注和常规自然语言处理技术,过程耗时费力,尤其在铸造制造等专业领域更为突出。大语言模型(LLMs)的兴起为自动化知识抽取提供了新可能。本研究比较了三种基于大模型的方法:预训练模型驱动法、上下文学习(ICL)法和微调法,旨在利用有限数据从领域文本中提取术语与关系。通过性能对比,选取最优方法构建铸造领域本体,并由领域专家进行验证。

原文摘要 · Abstract (English)

Ontologies are essential for structuring domain knowledge, improving accessibility, sharing, and reuse. However, traditional ontology construction relies on manual annotation and conventional natural language processing (NLP) techniques, making the process labour-intensive and costly, especially in specialised fields like casting manufacturing. The rise of Large Language Models (LLMs) offers new possibilities for automating knowledge extraction. This study investigates three LLM-based approaches, including pre-trained LLM-driven method, in-context learning (ICL) method and fine-tuning method to extract terms and relations from domain-specific texts using limited data. We compare their performances and use the best-performing method to build a casting ontology that validated by domian expert.

本体构建大模型应用信息抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。