用AI自动提取会议征稿信息,让非出版方学术活动可被系统研究
A Pathway for Assessing Grey Literature: Leveraging AI to Extract Conference Metadata and Organiser Information from Calls for Papers
- 构建多阶段AI流程,从征稿启事中抽取会议细节和组织者信息
- 实现会议届次、地点、组织者角色与机构的精准识别与去重
- 为分析非出版方学术会议提供可扩展的数据基础,适合科研政策研究者
尽管重要,灰色文献(如征稿启事)因格式不统一、高度异质,传统工具难以规模化处理,长期被元科学与科学计量学忽视。大型语言模型为此提供了新机遇。本文提出COCI框架,通过多阶段管道实现从原始征稿文本中自动化提取细粒度结构化元数据。该流程包含实体抽取、作者去重(基于OpenAlex)、主题与会议系列的语义映射,可准确识别会议届次、地理分布、组织者名单及其角色与隶属关系。通过结构化此前不可访问的信息,COCI为灰色文献的系统分析奠定基础,推动研究视角转向非出版方学术活动。
原文摘要 · Abstract (English)
Despite its importance, grey literature, including Calls for Papers (CfPs), remains largely overlooked in Metascience and Scientometric analysis due to its unstructured, highly heterogeneous format, which traditional tools struggle to process at scale. However, Large Language Models now offer a pivotal opportunity to devise innovative tools for systematically harvesting and processing such data. In this paper, we introduce COCI, an AI-based framework that automates the extraction of granular, structured metadata from raw CfP text. COCI employs a multi-stage pipeline for entity extraction, followed by author disambiguation against OpenAlex and semantic mapping of topics and conference series. This process identifies key data points, including conference editions, geographic locations, and comprehensive lists of organisers, along with their specific roles and affiliations. By structuring this previously inaccessible information, COCI establishes a foundation for the systematic analysis of grey literature, enabling new research opportunities and shifting the scholarly focus towards non-publisher-based events.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。