arXiv:2603.01092cs.AIcs.LG2026-03被引 2

用AI发现人类研究者想不到的科学方向,突破思维盲区。

The Alien Space of Science: Sampling Coherent but Cognitively Unavailable Research Directions

  • 将论文拆解为概念原子,构建可组合的科研词汇表
  • 同时评估方向的合理性与现有团队是否可能提出,筛选低可用性高合理性想法
  • 在1.6万篇论文上测试,探索范围扩大3.5-7倍,结果更优

科学发现不仅受限于事实,也受限于研究者当前的认知可用性。许多方向在现有文献框架下逻辑自洽,却因缺乏合适的概念、方法与直觉组合,难以被提出。现代语言模型继承了这种偏见,仅在文献高密度区域重组想法。本文提出一个新框架,聚焦“科学异域空间”——那些符合现有知识结构但未被现有研究社区覆盖的方向。方法首先将论文分解为细粒度概念单元,聚类形成统一的概念原子库;随后训练两个互补模型:共现模型评估组合是否构成可行研究方向,可用性模型评估是否有团队能提出该组合。采样时通过最大化共现性并最小化可用性,生成“异域”方向。在包含16,068篇来自NeurIPS、ICLR、ICML及主流NLP会议的论文语料上,该采样器比前沿大模型基线扩展了3.5至7倍的有效概念词汇量,且不牺牲合理性,在盲评大模型、人工评估和下游实验中表现相当或更优。该框架将科学可行性与社区可用性分离,使AI能补充而非复制人类科研,拓展到当前社区可能忽略的合理方向。

原文摘要 · Abstract (English)

Scientific discovery is constrained not only by what is true, but by what is cognitively available to the researchers currently exploring a field. Many directions are coherent in light of the literature yet unlikely to be proposed because no existing community occupies the right combination of concepts, methods, and intuitions. Modern language models inherit this bias, recombining high-density regions of the literature when prompted for novel ideas. We introduce a framework that targets the complementary region, which we call the alien space of science, where directions are plausible under the structure of existing knowledge but unlikely under the distribution of existing researchers. Our method first decomposes papers into granular conceptual units and clusters them into a shared vocabulary of idea atoms. It then learns two complementary models over this vocabulary. A coherence model scores whether a combination of atoms forms a viable research direction, and an availability model scores whether any existing author community is positioned to produce a given combination. Sampling alien directions then reduces to ranking atom combinations that maximize coherence while minimizing availability. On a corpus of 16,068 peer-reviewed LLM papers from NeurIPS, ICLR, ICML, and major NLP venues, the resulting sampler explores a 3.5 - 7 x broader effective atom vocabulary than frontier LLM ideation baselines without sacrificing coherence, and produces ideas that match or exceed those baselines under blind LLM, human, and downstream experimental evaluation. By separating scientific plausibility from community availability, our framework points toward AI ideation that complements rather than merely accelerates human science, expanding exploration into coherent directions that the current community may overlook.

科学发现AI创新研究方向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。