构建科学思想演化评估基准,测试AI能否理解研究的传承与创新关系。
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

- 用基因组框架表示论文思想,追踪继承、变异等六类演化行为。
- 14个大模型在推理任务中最高仅27.3%准确率,暴露组合能力瓶颈。
- 适合关注AI科研创造力、思想演化分析的研究者使用。
科学思想很少从零开始。它们继承机制、修复已有缺陷,并重组早期成果,如同生物基因组。现有基准仍无法衡量AI是否能捕捉这种继承结构。我们提出IdeaGene-Bench(IG-Bench),用于科学思想演化推理与基于谱系的思想生成评估。IG-Bench基于IdeaGene框架:每篇论文或提案被表示为一组最小化、类型化、证据支撑的思维基因组对象,通过GenomeDiff记录六种演化动态下的继承、突变、丢失、外部引入和新插入。基准包含1,961条真实谱系轨迹、1,085个精选思维基因组对象及920对GenomeDiff记录,覆盖10个科学领域。支持两项评估:IG-Exam(42种任务类型,1,029个实例)测试闭式谱系推理,涵盖抽象、溯源、演化推理与验证;IG-Arena采用谱系条件下的种群演化评分(PES),评估生成提案是否能作为合理后代:需继承正确基因组对象、与邻近工作有显著差异,并具备未来研究的选择价值。14个基于大模型的科研代理实验显示组合性瓶颈:最强系统在谱系推理中仅达27.3%精确率,结构化谱系上下文仅重排排名,未普遍提升所有模型表现。
原文摘要 · Abstract (English)
Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Bench), a benchmark for scientific lineage reasoning and lineage-grounded idea generation. IG-Bench is organized around the IdeaGene framework: each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns these objects to record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics. The benchmark contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records across 10 scientific domains. It supports two evaluations. IG-Exam (42 task types, 1,029 instances) tests closed-form lineage reasoning across Idea Genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification. IG-Arena evaluates generation with a lineage-conditioned Population-Evolution Score(PES), asking whether a proposal can be inserted as a coherent descendant of a given lineage population: it should inherit the right Idea Genome objects, vary meaningfully from nearby work, and offer selection value for future research. Experiments on 14 LLM-based scientists expose a compositional bottleneck. The strongest system reaches only 27.3% exact accuracy on lineage reasoning, and structured lineage context reshuffles system rankings rather than helping every participant uniformly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。