解决复杂机构名称匹配难题,提升科研数据质量
From raw affiliations to organization identifiers
- 采用先进解析与消歧技术处理多机构混杂的机构名
- 在复杂字符串中准确识别出目标机构标识符
- 适合需要高质量科研元数据的研究者与数据平台
准确的机构匹配对于提升科研元数据质量、推动全面文献计量分析以及支持学术知识库间的数据互操作至关重要。现有方法难以应对机构名称中常含多个机构或冗余信息的复杂情况。本文提出AffRo新方法,结合先进解析与消歧技术,有效应对上述挑战,并构建了专家标注的AffRoDB数据集,用于系统评估机构匹配算法,确保可靠基准测试。实验结果表明,AffRo能从复杂机构字符串中准确识别出组织标识符。
原文摘要 · Abstract (English)
Accurate affiliation matching, which links affiliation strings to standardized organization identifiers, is critical for improving research metadata quality, facilitating comprehensive bibliometric analyses, and supporting data interoperability across scholarly knowledge bases. Existing approaches fail to handle the complexity of affiliation strings that often include mentions of multiple organizations or extraneous information. In this paper, we present AffRo, a novel approach designed to address these challenges, leveraging advanced parsing and disambiguation techniques. We also introduce AffRoDB, an expert-curated dataset to systematically evaluate affiliation matching algorithms, ensuring robust benchmarking. Results demonstrate the effectiveness of AffRp in accurately identifying organizations from complex affiliation strings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。