arXiv:2605.05929cs.AI2026-05

首次为语义网知识图谱定义低资源语言分级标准

Which Are the Low-Resource Languages of the Semantic Web?

论文配图:Which Are the Low-Resource Languages of the Semantic Web?
图 1 · 摘自论文原文
  • 基于DBpedia等三大知识图谱分析语言分布,构建多层级分类体系
  • 提出低/中/高资源语言的量化划分标准,支持跨语言迁移选择
  • 为消除数字鸿沟提供可操作的语言资源评估框架,适合语义网研究者

新兴数字技术正加剧开放获取数据在高资源与低资源语言之间的不平等,使众多群体被排除在全球数字化进程之外。多语言链接开放数据知识图谱(LOD KGs)可通过跨语言迁移缓解这一问题,但目前尚无针对LOD KGs中低资源语言的明确量化定义。本文提出一种方法,分析语言在DBpedia、BabelNet和Wikidata中的分布情况,构建初步的多层级语言分类体系。该分类可用于正式定义低、中、高资源语言,未来可指导跨语言迁移候选语言的选择。

原文摘要 · Abstract (English)

Emerging digital technologies are exacerbating the existing divide in Open Access Data (OAD) between high-and low-resource languages, excluding many communities from the global digital transformation. Multilingual Linked Open Data Knowledge Graphs (LOD KGs) could contribute to mitigating this divide through cross-lingual transfer; however, no clear quantitative definition of low-resource languages has yet been established in the context of LOD KGs. In this poster, we present a methodology to analyze the distribution of languages across LOD KGs and propose a preliminary multi-level categorization based on DBpedia, BabelNet, and Wikidata. This categorization is leveraged to bring a formal definition of low-, high-, and medium-resource languages that could be later leveraged to select cross-lingual transfer candidates.

知识图谱多语言资源评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。