用范畴论结构挖掘可证伪的研究假说,提升自动科研创意质量。
Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure

- 将论文建模为带类型实体和关系的范畴,通过函子保结构匹配跨领域灵感。
- 在数万篇论文中筛选出17:1的候选方案,接受率超83%且可验证真伪。
- 不仅过滤无效假设,还保留每条拒绝理由,适合科研辅助与可解释创新。
基于大语言模型的自动化研究创意生成系统存在结构性缺陷:它们将创意生成简化为自由文本重组、随机论文配对或嵌入相似性检索。这些方法均将论文视为扁平对象(字符串或向量),忽略了研究人员在跨领域类比中实际使用的“问题-方法-度量-主张”关系链。本文引入范畴论中最基本的结构——复合与恒等箭头,恢复缺失的逻辑链条,使我们能够判断一个提议的类比是否保持关系连续性。具体地,每篇论文 $p$ 被建模为一个小范畴 $C_p$,其对象为提取出的类型化研究实体,态射为论文声明的关系;从 $p$ 到 $q$ 的跨论文桥梁是一个部分函子候选 $F: C_p \to C_q$,要求保持对象类型和覆盖的关系类别。我们实现了一个三层算法:范畴签名聚类、函子保真门控、六轴大模型可信度评估。在数万篇全文解析论文上,四种消融条件下,该结构门控实现了约17:1的候选过滤比,且接受方案的定量可证伪率始终高于83%;所有被拒方案均保留其各轴理由,使该系统兼具过滤与日志双重功能。
原文摘要 · Abstract (English)
Automated research-idea generation systems built on large language models (LLMs) share a structural weakness: they reduce ideation to free-text recombination, random paper pairing, or embedding-similarity retrieval. The three approaches fail in the same way: each treats a paper as a flat object, a string or a vector, and so quotients away the typed problem-method-metric-claim arrows a researcher actually uses when reasoning about a cross-domain analogy. We recover the missing structure with the minimal piece of category theory that a typed graph alone does not provide: composition, together with identity arrows, which makes it possible to ask whether a proposed analogy preserves relation chains. Concretely, each paper $p$ is modelled as a small category $C_p$ whose objects are extracted typed research entities and whose morphisms are the relations the paper asserts; a cross-paper bridge from $p$ to $q$ is then a partial functor candidate $F: C_p -> C_q$ that preserves object kinds and covered relation classes. We instantiate the model as a three-layer algorithm: categorical signature clustering, a functor-preservation gate, and a six-axis LLM plausibility judge. Evaluated on a corpus of tens of thousands of full-text-parsed papers under four ablation conditions, the categorical gate filters cross-domain candidates at roughly a 17:1 ratio while the quantitative-falsifier rate of accepted ideas stays above 83% throughout; every rejected candidate is retained with its per-axis rationale, so the gate doubles as a logging layer rather than a silent filter.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。