arXiv:2412.20005cs.CLcs.AI2024-12被引 20

基于多智能体的开源知识抽取系统,支持多领域文档自动提取

OneKE: A Dockerized Schema-Guided LLM Agent-based Knowledge Extraction System

  • 采用多智能体协同架构,按任务分工完成知识抽取
  • 在多个基准数据集上表现优异,支持灵活的模式配置与错误修正
  • 适合需要跨领域知识提取的研究者与开发者使用

我们提出OneKE,一个容器化部署的、基于模式引导的智能体知识抽取系统,可从网络文本和原始PDF书籍中提取知识,支持科学、新闻等多领域应用。系统设计了多个智能体协同工作,并配备可配置的知识库,实现不同任务场景下的高效抽取。知识库支持模式自定义、异常情况调试与修正,有效提升抽取性能。在基准数据集上的实证评估验证了其有效性,案例研究进一步展示了其在多领域任务中的适应性,凸显其广泛应用潜力。代码已开源(https://github.com/zjunlp/OneKE),演示视频可在http://oneke.openkg.cn/demo.mp4获取。

原文摘要 · Abstract (English)

We introduce OneKE, a dockerized schema-guided knowledge extraction system, which can extract knowledge from the Web and raw PDF Books, and support various domains (science, news, etc.). Specifically, we design OneKE with multiple agents and a configure knowledge base. Different agents perform their respective roles, enabling support for various extraction scenarios. The configure knowledge base facilitates schema configuration, error case debugging and correction, further improving the performance. Empirical evaluations on benchmark datasets demonstrate OneKE's efficacy, while case studies further elucidate its adaptability to diverse tasks across multiple domains, highlighting its potential for broad applications. We have open-sourced the Code at https://github.com/zjunlp/OneKE and released a Video at http://oneke.openkg.cn/demo.mp4.

知识抽取多智能体开放源码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。