arXiv:2603.00051cs.DLcs.AI2026-03KDD

用知识图谱构建文献任务评测工具,让大模型更懂专业领域。

LitBench: A Graph-Centric Large Language Model Benchmarking Tool For Literature Tasks

  • 基于领域文献构建知识子图,以图结构组织学术信息
  • 小模型在该工具上表现媲美GPT-4o等顶尖模型
  • 支持自定义领域,适合做文献智能分析的研究者

尽管大语言模型已成为文献相关任务的主流框架,但因难以关联知识并跨领域语境推理,仍无法胜任专业文献代理角色。为此,我们提出LitBench——一个面向文献任务的图中心化大模型评测工具。其核心是通过数据整理生成领域特定的文献子图,并基于节点与边的文本属性构建训练与评估数据集。该工具灵活支持任意领域(如高阶学科或交叉方向)的文献图谱构建。除数据集构建外,还定义了涵盖节点/边分析到相关工作生成等完整文献任务体系,使模型在训练中内化领域知识与关系,同时支持严谨评估。实验表明,使用LitBench数据集训练的小型领域模型性能可媲美GPT-4o和DeepSeek-R1。为提升可用性,我们开源该工具及配套AI代理,实现数据整理、模型训练与评估全流程自动化。

原文摘要 · Abstract (English)

While large language models (LLMs) have become the de facto framework for literature-related tasks, they still struggle to function as domain-specific literature agents due to their inability to connect pieces of knowledge and reason across domain-specific contexts, terminologies, and nomenclatures. This challenge underscores the need for a tool that facilitates such domain-specific adaptation and enables rigorous benchmarking across literature tasks. To that end, we introduce LitBench, a benchmarking tool designed to enable the development and evaluation of domain-specific LLMs tailored to literature-related tasks. At its core, LitBench uses a data curation process that generates domain-specific literature sub-graphs and constructs training and evaluation datasets based on the textual attributes of the resulting nodes and edges. The tool is designed for flexibility, supporting the curation of literature graphs across any domain chosen by the user, whether high-level fields or specialized interdisciplinary areas. In addition to dataset curation, LitBench defines a comprehensive suite of literature tasks, ranging from node and edge level analyses to advanced applications such as related work generation. These tasks enable LLMs to internalize domain-specific knowledge and relationships embedded in the curated graph during training, while also supporting rigorous evaluation of model performance. Our results show that small domain-specific LLMs trained and evaluated on LitBench datasets achieve competitive performance compared to state-of-the-art models like GPT-4o and DeepSeek-R1. To enhance accessibility and ease of use, we open-source the tool along with an AI agent tool that streamlines data curation, model training, and evaluation.

大模型评测知识图谱文献智能领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。