arXiv:2608.24559cs.DLcs.AI2026-08中稿 · ISWC 2026

让会议征稿信息自动变结构化数据,接入学术知识图谱。

COCI: Conference Organisers and Content Identifier

论文配图:COCI: Conference Organisers and Content Identifier
图 1 · 摘自论文原文
  • 用大模型+语义映射分步提取征稿文本中的实体信息。
  • 将作者、会议、主题等信息与OpenAlex等知识库对齐。
  • 适合关注学术活动数字化、知识图谱构建的研究者。

尽管灰文献在学术传播中扮演关键角色,但诸如征稿启事(CfPs)等资料仍与现代学术知识图谱严重脱节。这些文档结构松散且高度异质,长期阻碍其大规模处理。本文展示了一个名为会议组织者与内容标识器(COCI)的AI框架,旨在从原始征稿文本中提取细粒度、结构化的元数据。COCI采用多阶段流水线,结合大语言模型(LLMs)与语义映射技术,将提取出的实体与开放科学资源如OpenAlex、DBLP、TIB ConfIDent及AIDA Dashboard等知识库整合。通过消歧作者与语义对齐主题及会议系列,COCI弥合了非出版方学术活动的非正式传播与结构化语义网络之间的鸿沟,为非出版社主导的学术活动系统分析奠定基础。

原文摘要 · Abstract (English)

Despite the critical role of grey literature in scholarly communication, artefacts such as Calls for Papers (CfPs) remain largely isolated from modern Scholarly Knowledge Graphs. The unstructured and highly heterogeneous nature of these documents has traditionally hindered their large-scale processing. In this demo paper, we present the Conference Organisers and Content Identifier (COCI), an AI-based framework designed to extract fine-grained, structured metadata from raw CfP texts. COCI employs a multi-stage pipeline that combines Large Language Models (LLMs) with semantic mapping techniques to integrate extracted entities with established knowledge bases, including OpenAlex, DBLP, TIB ConfIDent, and the AIDA Dashboard. By disambiguating authors and semantically aligning topics and conference series, COCI bridges the gap between informal scholarly dissemination and structured Semantic Web resources, laying the foundation for systematic analysis of non-publisher-based academic events.

知识图谱信息抽取学术传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。