arXiv:2602.22215cs.AIcs.CL2026-02

用作者图谱+检索增强,让大模型生成有依据的科学创意。

Graph Your Way to Inspiration: Integrating Co-Author Graphs with Retrieval-Augmented Generation for Large Language Model Based Scientific Idea Generation

  • 构建作者知识图谱与灵感源采样,形成可追溯的外部知识库。
  • 融合RAG与GraphRAG,实现深度与广度兼顾的混合检索。
  • 通过强化学习优化提示词,提升创意的新颖性与相关性。

大型语言模型(LLMs)在科学创意生成中展现潜力,但生成结果常缺乏可控的学术背景和可追溯的灵感路径。本文提出GyWi系统,将作者知识图谱与检索增强生成(RAG)结合,构建外部知识库,为LLMs提供可控上下文与灵感溯源路径。首先提出以作者为中心的知识图谱构建方法及灵感源采样算法;其次设计融合RAG与GraphRAG的混合检索机制,获取兼具深度与广度的知识内容;第三,提出基于强化学习原理的提示词优化策略,自动引导LLM根据混合上下文优化输出。为评估方法,基于arXiv(2018–2023)构建评测数据集,并采用多维度评估方法:多项选择题自动评估、基于LLM的评分、人工评估及语义空间可视化分析。从新颖性、可行性、清晰度、相关性、重要性五个维度评估生成创意。在GPT-4o、DeepSeek-V3、Qwen3-8B、Gemini 2.5等主流模型上实验,结果表明,GYWI在新颖性、可靠性与相关性等指标上显著优于基线模型。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate potential in the field of scientific idea generation. However, the generated results often lack controllable academic context and traceable inspiration pathways. To bridge this gap, this paper proposes a scientific idea generation system called GYWI, which combines author knowledge graphs with retrieval-augmented generation (RAG) to form an external knowledge base to provide controllable context and trace of inspiration path for LLMs to generate new scientific ideas. We first propose an author-centered knowledge graph construction method and inspiration source sampling algorithms to construct external knowledge base. Then, we propose a hybrid retrieval mechanism that is composed of both RAG and GraphRAG to retrieve content with both depth and breadth knowledge. It forms a hybrid context. Thirdly, we propose a Prompt optimization strategy incorporating reinforcement learning principles to automatically guide LLMs optimizing the results based on the hybrid context. To evaluate the proposed approaches, we constructed an evaluation dataset based on arXiv (2018-2023). This paper also develops a comprehensive evaluation method including empirical automatic assessment in multiple-choice question task, LLM-based scoring, human evaluation, and semantic space visualization analysis. The generated ideas are evaluated from the following five dimensions: novelty, feasibility, clarity, relevance, and significance. We conducted experiments on different LLMs including GPT-4o, DeepSeek-V3, Qwen3-8B, and Gemini 2.5. Experimental results show that GYWI significantly outperforms mainstream LLMs in multiple metrics such as novelty, reliability, and relevance.

科学创意知识图谱检索增强大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。