arXiv:2608.07023cs.CLcs.AI2026-08

用大模型生成多语言技能知识图谱,自动识别和整合新技能。

An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation

  • 结合上下文与底层数据,动态构建技能节点和关系
  • 在五种欧洲语言中实现技能实体对齐与去重
  • 适合需要自更新技能库的招聘平台使用

人力资源平台面临数千条非标准化、多语言的专业能力声明组织难题,直接影响人才匹配等下游任务。为此,我们提出一种混合式知识图谱生成流程:将大语言模型(LLM)锚定在多语言维基数据(Wikidata)知识图谱上,同时采用代理反思机制合成新兴概念及其元数据。不同于僵化的自上而下或碎片化的自下而上方法,该系统将已知概念映射到稳定知识图实体,同时为未知技能动态创建新节点与关系元数据。流程包含五个阶段:实体对齐、多语言规范化、主动校正、去重及未映射概念的迭代恢复,在五种欧洲语言中自主适应快速演变、噪声密集的技能表述。最终生成一个可扩展、可解释且具备自修复能力的全面技能知识图谱,并从中提炼出结构化分类体系,全部基于非结构化、嘈杂文本。

原文摘要 · Abstract (English)

Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, directly impacting downstream tasks like accurate talent matching. To address this, we propose a hybrid knowledge graph generation pipeline that grounds a Large Language Model (LLM) in the Wikidata multilingual Knowledge Graph (KG) while employing an agentic reflexion pattern to synthesize emerging concepts and their associated metadata. Unlike rigid top-down methods or fragmented bottom-up approaches, our system anchors recognized concepts to stable Knowledge Graph entities while dynamically creating new nodes and relational metadata for unrecognized skills. Executed across five stages, entity reconciliation, multilingual canonicalization, active curation, deduplication, and the iterative recovery of unmapped concepts, the system autonomously adapts to rapidly evolving, noisy skill mentions across five European languages. Ultimately, this pipeline provides a highly scalable, explicable, and self-healing framework for generating a comprehensive skills knowledge graph, from which a structured taxonomy is derived, using unstructured, noisy text.

知识图谱多语言大模型技能挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。