重新设计正则化方法,让知识图谱补全模型突破性能瓶颈
Rethinking Regularization Methods for Knowledge Graph Completion
- 引入基于秩的稀疏正则化,有选择地惩罚重要特征分量
- 在多个数据集上显著降低过拟合,提升模型性能上限
- 适合追求高精度知识图谱补全的研究者和工程应用
知识图谱补全(KGC)近年来受到广泛关注,因其对提升知识图谱质量至关重要。尽管研究者持续探索各类模型,但以往工作多忽视从深层视角利用正则化,未能充分发挥其潜力。本文重新思考正则化在KGC中的应用。通过在多种KGC模型上的广泛实验,发现精心设计的正则化不仅能缓解过拟合、降低方差,还能使模型突破原有性能上限。为此,提出一种新型稀疏正则化方法SPR,将基于秩的选通稀疏性概念融入正则项,核心思想是选择性惩罚嵌入向量中具有显著特征的分量,从而有效忽略贡献小、可能仅代表噪声的成分。在多个数据集和模型上的对比实验表明,SPR优于其他正则化方法,能进一步突破性能边际。
原文摘要 · Abstract (English)
Knowledge graph completion (KGC) has attracted considerable attention in recent years because it is critical to improving the quality of knowledge graphs. Researchers have continuously explored various models. However, most previous efforts have neglected to take advantage of regularization from a deeper perspective and therefore have not been used to their full potential. This paper rethinks the application of regularization methods in KGC. Through extensive empirical studies on various KGC models, we find that carefully designed regularization not only alleviates overfitting and reduces variance but also enables these models to break through the upper bounds of their original performance. Furthermore, we introduce a novel sparse-regularization method that embeds the concept of rank-based selective sparsity into the KGC regularizer. The core idea is to selectively penalize those components with significant features in the embedding vector, thus effectively ignoring many components that contribute little and may only represent noise. Various comparative experiments on multiple datasets and multiple models show that the SPR regularization method is better than other regularization methods and can enable the KGC model to further break through the performance margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。