提升无监督实体对齐中伪种子质量,改善知识图谱覆盖均衡性
PSQE: A Theoretical-Practical Approach to Pseudo Seed Quality Enhancement for Unsupervised Multimodal Entity Alignment
- 利用多模态信息与聚类重采样增强伪种子精度与覆盖平衡
- 理论证明伪种子同时影响对比学习的吸引与排斥项
- 适用于提升各类无监督多模态实体对齐模型性能
多模态实体对齐(MMEA)旨在识别跨不同数据模态的等价实体,促进结构化数据融合,从而提升大语言模型应用表现。为降低对难获取标注种子对的要求,近期方法转向使用伪对齐种子的无监督范式。然而,多模态环境下无监督实体对齐仍研究不足,主要因多模态信息引入导致伪种子在知识图谱中覆盖不平衡。为此,本文提出PSQE(伪种子质量增强)方法,通过多模态信息与聚类重采样提升伪种子的精确性与图谱覆盖平衡性。理论分析揭示了伪种子对现有基于对比学习的MMEA模型的影响:伪种子可同时作用于对比学习中的吸引项与排斥项;而覆盖不均会导致模型优先关注高密度区域,削弱稀疏区域实体的学习能力。实验验证了理论发现,并表明PSQE作为即插即用模块,能显著提升基线模型性能。
原文摘要 · Abstract (English)
Multimodal Entity Alignment (MMEA) aims to identify equivalent entities across different data modalities, enabling structural data integration that in turn improves the performance of various large language model applications. To lift the requirement of labeled seed pairs that are difficult to obtain, recent methods shifted to an unsupervised paradigm using pseudo-alignment seeds. However, unsupervised entity alignment in multimodal settings remains underexplored, mainly because the incorporation of multimodal information often results in imbalanced coverage of pseudo-seeds within the knowledge graph. To overcome this, we propose PSQE (Pseudo-Seed Quality Enhancement) to improve the precision and graph coverage balance of pseudo seeds via multimodal information and clustering-resampling. Theoretical analysis reveals the impact of pseudo seeds on existing contrastive learning-based MMEA models. In particular, pseudo seeds can influence the attraction and the repulsion terms in contrastive learning at once, whereas imbalanced graph coverage causes models to prioritize high-density regions, thereby weakening their learning capability for entities in sparse regions. Experimental results validate our theoretical findings and show that PSQE as a plug-and-play module can improve the performance of baselines by considerable margins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。