构建首个科学创意重组知识库,助力跨领域研究灵感发现
CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation
- 从论文中自动挖掘创意重组案例,定义新信息抽取任务
- 覆盖人工智能与生物领域,支持跨学科研究方向分析
- 可用于训练生成式模型,提出被研究人员评价为有启发性的假说
人类创新的核心在于重组——将现有概念与机制融合以创造新想法。本文提出CHIMERA,首个大规模从科学文献中自动挖掘重组实例的知识库。该知识库支持对科学家如何跨领域汲取灵感进行实证分析,并可训练模型生成跨学科研究方向。我们定义了一项新的信息抽取任务:识别论文中的重组实例,构建专家标注数据集并微调基于大模型的抽取模型,应用于广泛的AI论文语料,同时验证其在生物学领域的泛化能力。通过两个应用展示其价值:一是分析人工智能各子领域中的重组模式;二是利用知识库训练科学假说生成模型,结果显示该模型提出的建议被研究人员评为具有启发性。
原文摘要 · Abstract (English)
A hallmark of human innovation is recombination -- the creation of novel ideas by integrating elements from existing concepts and mechanisms. In this work, we introduce CHIMERA, the first large-scale Knowledge Base (KB) of recombination examples automatically mined from the scientific literature. CHIMERA enables empirical analysis of how scientists recombine concepts and draw inspiration from different areas, and enables training models that propose cross-disciplinary research directions. To construct this KB, we define a new information extraction task: identifying recombination instances in papers. We curate an expert-annotated dataset and use it to fine-tune an LLM-based extraction model, which we apply to a broad corpus of AI papers. We also demonstrate generalization to a biological domain. We showcase the utility of CHIMERA through two applications. First, we analyze patterns of recombination across AI subfields. Second, we train a scientific hypothesis generation model using the KB, showing that it can propose directions that researchers rate as inspiring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。