利用多源不一致信息自动修正作者姓名歧义,提升学术搜索准确性。
Cross-Source Reasoning-based Correction for Author Name Disambiguation

- 通过多源数据不一致性检测,自动识别错误匹配
- 在真实数据集上比17个基线方法更优,无需人工标注
- 适合需要高精度作者识别的学术系统开发人员
作者姓名歧义是学术搜索系统中的关键挑战,现有方法常因论文-作者匹配的累积误差而失效,且忽略不同来源间的一致性问题。依赖专家标注成本过高。本文提出一种新视角:通过跨源不一致信息进行纠正。构建了全栈框架CrossND,包含数据清洗、跨源推理和测试时扩展三部分。首先,链式清洗流程去噪作者资料,提升论文-作者匹配概率准确性;其次,监督微调结合优化信号与基于概率软逻辑的跨源校正模块,推断出哪些来源的分配存在错误;最后,测试时扩展进一步增强预测精度与鲁棒性。在真实数据集上的实验表明,CrossND在无需人工干预的情况下,持续优于17个基线模型。
原文摘要 · Abstract (English)
Author name disambiguation is a critical challenge in academic search systems, often addressed through from-scratch and real-time disambiguation approaches. However, current algorithms remain vulnerable to cumulative errors of paper-author assignments and overlook inconsistent assignments across different sources. Resorting to expert annotation is resource-intensive. To this end, this paper explores a new perspective for author name disambiguation: cross-source correction by leveraging inconsistent assignments across sources. We propose CrossND, a full-stack framework that integrates data refinement, cross-source reasoning, and test-time scaling. First, a chain-of-refinement pipeline denoises author profiles and produces more accurate paper-author matching probabilities. Second, a supervised fine-tuning process incorporates these refined signals and a probabilistic soft logic-based cross-correction module to infer the assignments of which sources are incorrect. Third, test-time scaling further enhances the accuracy and robustness of the predictions. Experiments on real-world datasets indicate that CrossND consistently outperforms 17 baselines by leveraging cross-source reasoning without human intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。