通过投票融合多类本体对齐方法,提升精度与召回平衡。
OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques

- 采用两阶段投票融合策略整合不同对齐方法的预测结果。
- 异构方法组合提升精度,同质LLM组合更优整体F1分数。
- 适用于需兼顾准确率与覆盖率的跨领域本体对齐任务。
本体对齐(OA)已历经从词法、结构对齐到知识图谱嵌入(KGE)模型,以及近期基于大语言模型(LLM)的方法。尽管现代OA框架为部署多种异构对齐器提供了统一生态,但系统性整合其互补且时常冲突的预测结果机制仍相对薄弱。我们提出OntoAligner-Ensemble,一个模块化、对齐器无关的框架,通过可配置的两阶段流程——基于投票的融合策略与后融合选择策略——整合候选对应关系。该框架支持任何在OntoAligner中实现并生成候选对应关系的对齐器,使不同对齐范式可通过统一决策过程集成。为验证有效性,我们使用代表性轻量级字符串对齐器、基于KGE的对齐器及依托开源权重与API调用式LLM的检索增强生成对齐器实例化框架。在涵盖生物医学到非等价领域的五个OAEI赛道中的八个基准任务上评估了独立对齐器与集成配置。结果表明,集成融合持续改善精确率与召回率之间的平衡,并在多数跨领域场景下优于单一对齐器。此外分析显示,集成组成直接影响精确率-召回率权衡:异构跨范式集成通常提升精确率,而同质LLM集成则更常获得更高整体F1分数。这些发现表明,系统性集成学习为本体对齐提供了稳健且可复现的策略,并为不同对齐场景下的集成组合选择提供实践指导。
原文摘要 · Abstract (English)
Ontology alignment (OA) has evolved through several methodological paradigms, ranging from lexical and structural aligners to knowledge graph embedding (KGE) models and, more recently, Large Language Model (LLM)-based approaches. Although modern OA frameworks provide unified ecosystems for deploying these heterogeneous aligners, mechanisms for systematically reconciling their complementary and sometimes conflicting predictions remain relatively underexplored. We present OntoAligner-Ensemble, a modular and aligner-agnostic framework that combines candidate correspondences through a configurable two-stage process comprising voting-based fusion strategies followed by post-fusion selection policies. The framework supports any aligner implemented within OntoAligner that produces candidate correspondences, enabling diverse alignment paradigms to be integrated through a unified decision process. To demonstrate its effectiveness, we instantiate the framework using representative lightweight string-aligner, KGE-based, and Retrieval-Augmented Generation aligners powered by both open-weight and API-based LLMs. We evaluate individual aligners and ensemble configurations across eight benchmark tasks from five OAEI tracks spanning biomedical to beyond-equivalence. The results show that ensemble fusion consistently improves the balance between precision and recall and frequently outperforms standalone aligners across diverse domains. Furthermore, our analysis reveals that ensemble composition directly affects the precision-recall trade-off: heterogeneous cross-paradigm ensembles generally improve precision, whereas homogeneous LLM ensembles more often achieve higher overall F1-scores. These findings demonstrate that systematic ensemble learning offers a robust and reproducible strategy for OA while providing practical guidance for selecting ensemble compositions under different alignment scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。