用大模型自动填充学术会议数据,让维基知识库更完整。
Scholarly Wikidata: Population and Exploration of Conference Data in Wikidata using LLMs
- 用大模型从官网和论文集提取会议元信息,仅需少量人工校验。
- 一次性为维基数据新增6000多个学术实体,覆盖105个会议。
- 方法可推广至其他领域,适合想共建开放学术资源的研究者。
已有多个项目尝试用本体建模学术数据并构建知识图谱,但自动化填充机制仍缺失,且语义网社区的各类倡议彼此孤立。本文提出利用维基数据基础设施,通过大模型从非结构化来源(如会议网站、论文集)及现有结构化数据中自动填充学术信息,实现可持续的数据更新。初步分析显示,语义网相关会议在维基数据中代表性不足。本研究主要贡献包括:(a) 分析现有学术数据本体,识别维基数据中的缺失实体与属性;(b) 基于大模型实现半自动提取,获取会议接受率、组织角色、程序委员会成员、最佳论文奖、主题报告和赞助商等信息,仅需最小程度人工验证;(c) 扩展维基数据可视化工具,支持生成数据的探索。研究聚焦105个语义网相关会议,共扩展/新增超过6000个实体。该方法具有普适性,可应用于其他学术领域,提升维基数据作为综合性学术资源的实用性。
原文摘要 · Abstract (English)
Several initiatives have been undertaken to conceptually model the domain of scholarly data using ontologies and to create respective Knowledge Graphs. Yet, the full potential seems unleashed, as automated means for automatic population of said ontologies are lacking, and respective initiatives from the Semantic Web community are not necessarily connected: we propose to make scholarly data more sustainably accessible by leveraging Wikidata's infrastructure and automating its population in a sustainable manner through LLMs by tapping into unstructured sources like conference Web sites and proceedings texts as well as already existing structured conference datasets. While an initial analysis shows that Semantic Web conferences are only minimally represented in Wikidata, we argue that our methodology can help to populate, evolve and maintain scholarly data as a community within Wikidata. Our main contributions include (a) an analysis of ontologies for representing scholarly data to identify gaps and relevant entities/properties in Wikidata, (b) semi-automated extraction -- requiring (minimal) manual validation -- of conference metadata (e.g., acceptance rates, organizer roles, programme committee members, best paper awards, keynotes, and sponsors) from websites and proceedings texts using LLMs. Finally, we discuss (c) extensions to visualization tools in the Wikidata context for data exploration of the generated scholarly data. Our study focuses on data from 105 Semantic Web-related conferences and extends/adds more than 6000 entities in Wikidata. It is important to note that the method can be more generally applicable beyond Semantic Web-related conferences for enhancing Wikidata's utility as a comprehensive scholarly resource. Source Repository: https://github.com/scholarly-wikidata/ DOI: https://doi.org/10.5281/zenodo.10989709 License: Creative Commons CC0 (Data), MIT (Code)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。