综述2016-2024年深度学习在作者消歧中的进展与挑战
Recent Developments in Deep Learning-based Author Name Disambiguation
- 系统梳理深度学习在作者消歧中的最新方法
- 混合模型有效融合结构化与非结构化数据
- 适合研究数字图书馆、信息检索的学者参考
作者姓名消歧(AND)是数字图书馆将作者与其发表文献正确关联的关键任务。由于研究人员缺乏持久标识符以及同名异义等语言挑战,深度学习算法在此问题上的应用日益广泛。尽管已有多种深度学习方法被提出,并存在对比技术、复杂度与性能的综述,但尚未有研究专门针对2016至2024年间基于深度学习的最新进展进行系统性回顾。本文对这一时期先进的深度学习型作者消歧技术进行了系统性综述,重点分析了近期改进、现存挑战与未解决问题。研究表明,深度学习方法显著推动了AND的发展,使其能够有效整合结构化与非结构化数据,而混合方法则在监督与无监督学习之间实现了良好平衡。
原文摘要 · Abstract (English)
Author Name Disambiguation (AND) is a critical task for digital libraries aiming to link existing authors with their respective publications. Due to the lack of persistent identifiers used by researchers and the presence of intrinsic linguistic challenges, such as homonymy, the development of Deep Learning algorithms to address this issue has become widespread. Many AND deep learning methods have been developed, and surveys exist comparing the approaches in terms of techniques, complexity, performance. However, none explicitly addresses AND methods in the context of deep learning in the latest years (i.e. timeframe 2016-2024). In this paper, we provide a systematic review of state-of-the-art AND techniques based on deep learning, highlighting recent improvements, challenges, and open issues in the field. We find that DL methods have significantly impacted AND by enabling the integration of structured and unstructured data, and hybrid approaches effectively balance supervised and unsupervised learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。