梳理维基数据限定词用法,构建实用分类体系
Understanding Wikidata Qualifiers: An Analysis and Taxonomy
- 基于频率与多样性分析,用改进熵指数筛选关键限定词
- 提炼出300个核心限定词,分为上下文、不确定性等四类
- 助力知识图谱构建与智能推荐,适合数据治理与语义研究者
本文深入分析维基数据限定词的语义与实际使用情况,旨在构建能解决限定词选择、图查询与逻辑推理挑战的分类体系。通过分析维基数据快照,采用改进的香农熵指数评估限定词重要性,以应对长尾现象。研究筛选出前300个高频且多样化的限定词,并将其归类为上下文型、认知/不确定性型、结构型及其他型限定词。该分类体系有助于指导贡献者创建和查询数据,提升限定词推荐系统性能,优化知识图谱设计方法。结果表明,该分类有效覆盖了最重要限定词,提供了结构化理解与利用限定词的框架。
原文摘要 · Abstract (English)
This paper presents an in-depth analysis of Wikidata qualifiers, focusing on their semantics and actual usage, with the aim of developing a taxonomy that addresses the challenges of selecting appropriate qualifiers, querying the graph, and making logical inferences. The study evaluates qualifier importance based on frequency and diversity, using a modified Shannon entropy index to account for the "long tail" phenomenon. By analyzing a Wikidata dump, the top 300 qualifiers were selected and categorized into a refined taxonomy that includes contextual, epistemic/uncertainty, structural, and additional qualifiers. The taxonomy aims to guide contributors in creating and querying statements, improve qualifier recommendation systems, and enhance knowledge graph design methodologies. The results show that the taxonomy effectively covers the most important qualifiers and provides a structured approach to understanding and utilizing qualifiers in Wikidata.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。