arXiv:2502.11610cs.IR2025-02被引 1

对比OpenAlex与Clarivate学者ID在不同领域的识别准确率。

Accuracy Assessment of OpenAlex and Clarivate Scholar ID with an LLM-Assisted Benchmark

  • 用LLM辅助标注100人×4组作者论文,构建基准数据集。
  • 跨领域测试显示,两者在高产作者中召回率均超85%。
  • 适合做科研影响力分析的学者身份匹配,尤其关注多产研究者。

在定量科学学研究中,精准识别学者对科学数据分析至关重要。然而,由于姓名常见、缩写及拼写习惯差异,该任务极具挑战。尽管ORCID等标识系统正在发展,但许多学者未注册,且大量文献未被收录。学术数据库如Clarivate和OpenAlex已推出自有标识系统,作为初步的姓名消歧方案。本研究评估了这些系统在不同群体中的有效性,以确定其适用场景。我们根据国家、学科和通讯作者论文数量,从Web of Science(WOS)期刊的前四分位(Q1)中选取作者,每组100人,利用增强搜索的大语言模型方法,对每位作者的所有论文进行细致标注。基于标注结果,分别在OpenAlex和Clarivate中查找对应标识,提取所有相关论文,筛选出发表于Q1 WOS期刊的论文,并通过与标注数据集对比,计算精确率与召回率。

原文摘要 · Abstract (English)

In quantitative SciSci (science of science) studies, accurately identifying individual scholars is paramount for scientific data analysis. However, the variability in how names are represented-due to commonality, abbreviations, and different spelling conventions-complicates this task. While identifier systems like ORCID are being developed, many scholars remain unregistered, and numerous publications are not included. Scholarly databases such as Clarivate and OpenAlex have introduced their own ID systems as preliminary name disambiguation solutions. This study evaluates the effectiveness of these systems across different groups to determine their suitability for various application scenarios. We sampled authors from the top quartile (Q1) of Web of Science (WOS) journals based on country, discipline, and number of corresponding author papers. For each group, we selected 100 scholars and meticulously annotated all their papers using a Search-enhanced Large Language Model method. Using these annotations, we identified the corresponding IDs in OpenAlex and Clarivate, extracted all associated papers, filtered for Q1 WOS journals, and calculated precision and recall by comparing against the annotated dataset.

学者识别数据质量命名消歧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。