用多语言搜索增强大模型,提升学者姓名消歧准确率
Scholar Name Disambiguation with Search-enhanced LLM Across Language
- 结合搜索引擎与多语言大模型,自动获取跨语种学术信息
- 在多语言数据上显著提升消歧效果,尤其对非英语学者
- 适合需要精准学术身份识别的评审、反欺诈等场景
学者姓名消歧在奖项候选人评估、申请材料防伪等实际场景中至关重要。尽管已有进展,现有方法仍受限于异构数据复杂性,常需大量人工干预。本文提出一种基于多语言搜索增强大模型的新方法,利用搜索引擎的查询重写、意图识别与数据索引能力,获取更丰富的实体信息与个人资料,拓展数据维度。得益于大语言模型强大的跨语言能力,结合优化的检索技术,可实现高效信息获取与利用。实验表明,引入本地语言显著提升消歧性能,尤其在地理分布多样学者群体中表现突出。该多语言、搜索增强的方法为高效、精准的主动式学者姓名消歧提供了新方向。
原文摘要 · Abstract (English)
The task of scholar name disambiguation is crucial in various real-world scenarios, including bibliometric-based candidate evaluation for awards, application material anti-fraud measures, and more. Despite significant advancements, current methods face limitations due to the complexity of heterogeneous data, often necessitating extensive human intervention. This paper proposes a novel approach by leveraging search-enhanced language models across multiple languages to improve name disambiguation. By utilizing the powerful query rewriting, intent recognition, and data indexing capabilities of search engines, our method can gather richer information for distinguishing between entities and extracting profiles, resulting in a more comprehensive data dimension. Given the strong cross-language capabilities of large language models(LLMs), optimizing enhanced retrieval methods with this technology offers substantial potential for high-efficiency information retrieval and utilization. Our experiments demonstrate that incorporating local languages significantly enhances disambiguation performance, particularly for scholars from diverse geographic regions. This multi-lingual, search-enhanced methodology offers a promising direction for more efficient and accurate active scholar name disambiguation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。