用引用演化图监督大模型,生成更连贯的科研创意。
Graphs of Research: Citation Evolution Graphs as Supervision for Research Idea Generation

- 构建论文演化有向无环图,融合引用位置、频率等关系信号。
- 在50组测试数据上,模型生成创意质量超越GPT-4o基线。
- 适合想提升科研创意生成能力的研究者使用。
科研创意生成是自动化科学研究的核心驱动力。尽管大语言模型(LLMs)在大规模创意生成方面展现出潜力,但现有方法主要依赖静态文献检索或复杂提示工程,未能利用参考文献间的结构关系。本文提出「研究图谱」(GoR),通过提取每篇种子论文的2跳引用邻域,结合引用位置、频率、前驱链接和发表时间,构建论文演化有向无环图(DAG)。我们建立自动化抽取流水线,从五大机器学习与自然语言处理顶会收集数据,包含498/50/50组训练/验证/测试种子论文及约7,600条被引参考文献。使用Qwen2.5-7B-Instruct-1M模型,在包含引用图、边信号、参考信息与任务定义的结构化提示下进行微调,以预测种子论文的科研创意。在与GPT-4o驱动基线的对抗性评测中,GoR-SFT表现最优,证明引用演化图作为监督信号的有效性。我们希望降低引用演化图的使用门槛,加速自动化科学创新。
原文摘要 · Abstract (English)
Research idea generation is the innovation-driving step of automated scientific research. Recently, large language models (LLMs) have shown potential for automating idea generation at scale. However, existing methods mainly condition LLMs on eliciting idea generation through static retrieval of relevant literature or complex prompt engineering, without discarding the structural relations among references. We propose Graphs of Research (GoR), a supervised fine-tuning method that extracts a 2-hop reference neighborhood for each seed paper, derives the relations among those references from citation position, frequency, predecessor links, and publication time, and organizes them into a paper-evolution directed acyclic graph (DAG). We construct an automated extraction pipeline that draws data from five major ML/NLP venues, comprising 498/50/50 train/validation/test seed papers and approximately 7,600 cited references. Qwen2.5-7B-Instruct-1M is fine-tuned on a structured-text prompt that includes the citation graph, edge signals, reference information, and task definition to predict the idea for the seed paper. Across head-to-head LLM-judge tournaments against gpt-4o-driven baselines, GoR-SFT achieves SOTA, demonstrating the effectiveness of citation-evolution graphs as supervision signal for LLM-based idea generation. We hope that this reduces the barrier for citation evolution graphs as a supervision, accelerating automated scientific innovation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。