用合成数据微调大模型,提升知识图谱对齐效果
Improving LLM-based Ontology Matching with fine-tuning on synthetic data
- 自动筛选相关模块生成提示,降低匹配复杂度
- 合成数据集使零样本性能提升显著,最高达15.6%准确率
- 适合知识工程、语义网研究者快速部署对齐工具
大型语言模型(LLMs)正被越来越多地应用于本体匹配流程中。本文研究了LLM直接在本体模块上执行匹配并生成对应对齐的能力,并探索了专用微调策略如何提升其在零样本设置下的表现。所提方法结合搜索空间缩减技术,从源与目标本体中选取相关子集,自动生成提示。针对训练所需参考对齐数据稀缺的问题,提出一种基于LLM的新型合成数据生成方法,构建了包含本体子模块对及其参考对齐的语料库,专用于微调LLM以完成本体匹配任务。该方法在OAEI复杂赛道的Conference、Geolink、Enslaved、Taxon和Hydrography数据集上进行了评估。结果表明,经过合成数据微调的LLM在零样本场景下表现优于未微调的基础模型。关键贡献在于将自动数据生成与微调相结合,有效适配LLM用于本体匹配任务。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly being integrated into various components of Ontology Matching pipelines. This paper investigates the capability of LLMs to perform ontology matching directly on ontology modules and generate the corresponding alignments. Furthermore, it is explored how a dedicated fine-tuning strategy can enhance the model's matching performance in a zero-shot setting. The proposed method incorporates a search space reduction technique to select relevant subsets from both source and target ontologies, which are then used to automatically construct prompts. Recognizing the scarcity of reference alignments for training, a novel LLM-based approach is introduced for generating a synthetic dataset. This process creates a corpus of ontology submodule pairs and their corresponding reference alignments, specifically designed to fine-tune an LLM for the ontology matching task. The proposed approach was evaluated on the Conference, Geolink, Enslaved, Taxon, and Hydrography datasets from the OAEI complex track. The results demonstrate that the LLM fine-tuned on the synthetically generated data exhibits superior performance compared to the non-fine-tuned base model. The key contribution is a strategy that combines automatic dataset generation with fine-tuning to effectively adapt LLMs for ontology matching tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。