用知识图谱提升低资源语言的数字可见性
In Data or Invisible: Toward a Better Digital Representation of Low-Resource Languages with Knowledge Graphs

- 分析三大知识图谱中语言分布差异,定位低资源语言缺失问题
- 研究跨语言迁移中基于语言相似性的候选选择策略
- 探索类比推理提升多语言知识图谱补全,扩大低语种覆盖
新兴数字技术正加剧高资源与低资源语言在开放数据上的差距,使众多社区难以参与全球数字化进程。本博士提案聚焦于链接开放数据知识图谱(LOD KGs)的语言覆盖问题。首先,通过分析三个主要多语言知识图谱——DBpedia、BabelNet 和 Wikidata——中的关键变量,包括各语言版本维基百科文章数及知识图谱中标注语言的实体数量,揭示语言在 LOD 中的分布特征。在此基础上,拟研究跨语言迁移中候选选择策略对多语言知识图谱补全任务的影响,重点考察基于语言相近性及已有标注对齐数据的策略。语言相似性也启发我们探索尚未被充分研究的类比推理机制,利用语言间的相似或相异关系识别对应实体,从而提升知识图谱补全性能并增强低资源语言的覆盖范围。
原文摘要 · Abstract (English)
Emerging digital technologies are exacerbating the existing divide in Open Access Data (OAD) between high-and low-resource languages, excluding many communities from participating in the global digital transformation. In this PhD proposal, we aim to address this gap, focusing on the language coverage of Linked Open Data knowledge graphs (LOD KGs). First, we identify key variables that characterize language distribution in LOD, including the number of Wikipedia articles per language edition and the number of language-tagged entities in LOD KGs. These variables are analyzed across three major multilingual LOD KGs, DBpedia, BabelNet, and Wikidata, providing insights into the representation and distribution of languages within LOD. Building on this analysis, we intend to study the impact of cross-lingual transfer candidate selection on the task of multilingual KG completion. In particular, we plan to investigate strategies based on linguistic proximity and the availability of curated annotated alignments between languages. Language proximity also motivates us to explore the benefits of analogical reasoning that relies on (dis)similarities and has not yet been investigated to identify correspondences across languages to improve KG completion performance and enhance language coverage in LOD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。