用模式引导的上下文学习实现多知识库复杂对齐,提升生物医学数据整合精度。
CMOMgen: Complex Multi-Ontology Alignment via Pattern-Guided In-Context Learning
- 通过检索增强生成选择相关类并构建复合映射,实现端到端对齐。
- 在三个生物医学任务中F1最低达63%,优于所有基线和消融版本。
- 46%的非参考映射获最高评分,证明其映射语义合理性强,适合领域整合场景。
构建全面的知识图谱需整合多个本体以充分上下文化数据。本体匹配通过发现跨本体概念等价关系,建立统一语义层。尽管现有简单成对匹配技术成熟,但单一等价映射无法实现相关但分离本体的完整语义整合。复杂多本体匹配(CMOM)将源实体对齐至多个目标实体的复合逻辑表达式,建立更精细的等价关系及本体层级溯源。本文提出首个端到端的CMOM策略——CMOMgen,可生成完整且语义正确的映射,不限制目标本体或实体数量。检索增强生成选择相关类以构造映射,并筛选匹配参考映射作为上下文示例,强化上下文学习效果。在三个生物医学任务中,以部分参考对齐为基准进行评估。CMOMgen在类选择上优于基线,验证了专用策略的有效性。其在两个任务中达到最高F1(最低63%),第三任务位列第二。此外,人工评估显示46%的非参考映射获得最高评分,进一步证实其构建语义合理映射的能力。
原文摘要 · Abstract (English)
Constructing comprehensive knowledge graphs requires the use of multiple ontologies in order to fully contextualize data into a domain. Ontology matching finds equivalences between concepts interconnecting ontologies and creating a cohesive semantic layer. While the simple pairwise state of the art is well established, simple equivalence mappings cannot provide full semantic integration of related but disjoint ontologies. Complex multi-ontology matching (CMOM) aligns one source entity to composite logical expressions of multiple target entities, establishing more nuanced equivalences and provenance along the ontological hierarchy. We present CMOMgen, the first end-to-end CMOM strategy that generates complete and semantically sound mappings, without establishing any restrictions on the number of target ontologies or entities. Retrieval-Augmented Generation selects relevant classes to compose the mapping and filters matching reference mappings to serve as examples, enhancing In-Context Learning. The strategy was evaluated in three biomedical tasks with partial reference alignments. CMOMgen outperforms baselines in class selection, demonstrating the impact of having a dedicated strategy. Our strategy also achieves a minimum of 63% in F1-score, outperforming all baselines and ablated versions in two out of three tasks and placing second in the third. Furthermore, a manual evaluation of non-reference mappings showed that 46% of the mappings achieve the maximum score, further substantiating its ability to construct semantically sound mappings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。