arXiv:2511.04926cs.CL2025-11被引 2

发现并修复维基数据分类体系中的语义不一致问题

Diagnosing and Mitigating Semantic Inconsistencies in Wikidata's Classification Hierarchy

  • 提出新验证方法检测分类错误和冗余关系
  • 发现特定领域存在大量过泛子类链接问题
  • 提供交互工具让用户审查任意实体的分类关系

Wikidata 是目前全球最大的开放知识图谱,包含超过1.2亿个实体。它整合了多个领域数据库的数据,并从维基百科导入大量内容,同时允许用户自由编辑。这种开放性使其成为知识图谱研究的核心资源,但也导致分类体系存在一定程度的不一致。本研究基于前期工作,提出并应用一种新的验证方法,确认了特定领域中存在分类错误、过度泛化的子类链接及冗余连接。我们进一步提出了判断问题是否需修正的新评估标准,并开发了一个系统,使用户可检查任意 Wikidata 实体的分类关系,充分发挥平台的众包优势。

原文摘要 · Abstract (English)

Wikidata is currently the largest open knowledge graph on the web, encompassing over 120 million entities. It integrates data from various domain-specific databases and imports a substantial amount of content from Wikipedia, while also allowing users to freely edit its content. This openness has positioned Wikidata as a central resource in knowledge graph research and has enabled convenient knowledge access for users worldwide. However, its relatively loose editorial policy has also led to a degree of taxonomic inconsistency. Building on prior work, this study proposes and applies a novel validation method to confirm the presence of classification errors, over-generalized subclass links, and redundant connections in specific domains of Wikidata. We further introduce a new evaluation criterion for determining whether such issues warrant correction and develop a system that allows users to inspect the taxonomic relationships of arbitrary Wikidata entities-leveraging the platform's crowdsourced nature to its full potential.

知识图谱数据清洗维基数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。