用论文引用关系找科研灵感,提升发现新方法的效率。
MIR: Methodology Inspiration Retrieval for Scientific Research Problems
- 构建方法传承图,通过引用关系捕捉科研思路演化
- 召回率提升5.4,平均精度提升7.8,超越强基线
- 适合需要跨领域启发的科研人员和AI辅助创新研究
大型语言模型在加速科学发现方面备受关注。现有方法依赖文献检索,但效果受文献质量与内容影响大。本文提出方法论灵感检索(MIR),旨在找到能为当前研究问题提供灵感的先前工作。为此,我们构建了专用于训练和评估的MIR数据集,并建立基线。提出方法论邻接图(MAG),通过引用关系刻画方法传承路径,将“直观先验”融入密集检索器,识别超越表面语义相似性的方法启发模式。实验显示,相比强基线,召回率提升+5.4(Recall@3),平均精度提升+7.8(mAP)。进一步结合基于LLM的重排序策略,再提升+4.5(Recall@3)和+4.8(mAP)。通过大量消融实验与定性分析,验证了MIR在推动自动化科学发现中的潜力,并指明灵感驱动检索的发展方向。
原文摘要 · Abstract (English)
There has been a surge of interest in harnessing the reasoning capabilities of Large Language Models (LLMs) to accelerate scientific discovery. While existing approaches rely on grounding the discovery process within the relevant literature, effectiveness varies significantly with the quality and nature of the retrieved literature. We address the challenge of retrieving prior work whose concepts can inspire solutions for a given research problem, a task we define as Methodology Inspiration Retrieval (MIR). We construct a novel dataset tailored for training and evaluating retrievers on MIR, and establish baselines. To address MIR, we build the Methodology Adjacency Graph (MAG); capturing methodological lineage through citation relationships. We leverage MAG to embed an "intuitive prior" into dense retrievers for identifying patterns of methodological inspiration beyond superficial semantic similarity. This achieves significant gains of +5.4 in Recall@3 and +7.8 in Mean Average Precision (mAP) over strong baselines. Further, we adapt LLM-based re-ranking strategies to MIR, yielding additional improvements of +4.5 in Recall@3 and +4.8 in mAP. Through extensive ablation studies and qualitative analyses, we exhibit the promise of MIR in enhancing automated scientific discovery and outline avenues for advancing inspiration-driven retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。