通过动态演化技能生成高质量知识超图,解决跨领域术语泛化难题
Hyper-KGGen: A Skill-Driven Knowledge Extractor for High-Quality Knowledge Hypergraph Generation
- 将抽取过程设计为动态技能进化,逐步分解文档覆盖多维关系
- 引入稳定性反馈机制,从不稳定的预测中提炼高质量领域技能
- 在多场景下优于静态少样本提示,适合复杂知识建模任务
知识超图通过封装复杂的n元原子事实,超越传统二元知识图谱,提供更全面的语义表示范式。然而,构建高质量超图仍具挑战,主要源于场景差距:通用抽取器难以在具有特定术语的多样化领域间泛化,而现有方法常无法平衡结构骨架与细粒度细节。为此,我们提出Hyper-KGGen,一种以技能驱动的框架,将抽取重构为动态技能演化过程。首先,采用粗粒度到细粒度机制系统分解文档,确保从二元链接到复杂超边的全维度覆盖。关键在于,其引入自适应技能获取模块,通过稳定性反馈回路将领域专长主动提炼至全局技能库中,以抽取稳定性作为相对奖励信号,从不稳定轨迹和遗漏预测中诱导出高质量技能。此外,我们构建了HyperDocRED,一个严格标注的文档级知识超图抽取基准。实验表明,Hyper-KGGen显著优于强基线,在多场景设置下,演化技能提供的指导远胜于静态少样本示例。
原文摘要 · Abstract (English)
Knowledge hypergraphs surpass traditional binary knowledge graphs by encapsulating complex n-ary atomic facts, providing a more comprehensive paradigm for semantic representation. However, constructing high-quality hypergraphs remains challenging due to the scenario gap: generic extractors struggle to generalize across diverse domains with specific jargon, while existing methods often fail to balance structural skeletons with fine-grained details. To bridge this gap, we propose Hyper-KGGen, a skill-driven framework that reformulates extraction as a dynamic skill-evolving process. First, Hyper-KGGen employs a coarse-to-fine mechanism to systematically decompose documents, ensuring full-dimensional coverage from binary links to complex hyperedges. Crucially, it incorporates an adaptive skill acquisition module that actively distills domain expertise into a Global Skill Library. This is achieved via a stability-based feedback loop, where extraction stability serves as a relative reward signal to induce high-quality skills from unstable traces and missed predictions. Additionally, we present HyperDocRED, a rigorously annotated benchmark for document-level knowledge hypergraph extraction. Experiments demonstrate that Hyper-KGGen significantly outperforms strong baselines, validating that evolved skills provide substantially richer guidance than static few-shot examples in multi-scenario settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。