arXiv:2604.12243cs.CLcs.AI2026-04被引 1

让AI持续追踪文献,提前预测未来研究热点并验证假设。

Continuous Knowledge Metabolism: Generating Scientific Hypotheses from Evolving Literature

  • 用滑动窗口动态更新知识,实现文献的持续学习
  • 生成的假设平均领先发表404天,36个主题中有55次命中未来论文
  • 适合关注前沿趋势的科研人员和需要预判方向的团队

在快速发展的子领域中识别有潜力的研究方向是现代人工智能研究中最耗费认知资源的任务之一。现有基于大模型的科学发现系统通常仅对静态文献快照进行一次性提示,且仅通过当前评审(如人工审稿、代理同行评审、实验验证或自评估)进行验证,无法判断其是否能预见未来趋势。我们提出连续知识代谢(CKM),一种用于假设生成的AI工作流,具备三大核心能力:(i) 通过滑动窗口实现持续的文献代谢,保持动态知识状态;(ii) 预测性评估,即根据生成窗口之后发表的论文来评判假设;(iii) 实践级故障检测,从输出中诊断工作流失败模式。在50个机器学习主题的基准测试中,CKM-Lite在72%的主题上至少产生一个被验证的假设(50个主题中的36个),相比单次提示基线(30%)提升超过一倍,每主题成本约3美元,令牌消耗降低91%。被验证的假设平均比匹配论文提前404天(36个主题中共55次命中,中位数399天,范围66–757天)。总体而言,对未来发展文献的预测性验证为可证伪、低成本的评估协议提供了替代方案,适用于所有具有发布时间记录的语料库。

原文摘要 · Abstract (English)

Identifying promising research directions in fast-moving subareas is one of the most cognitively expensive tasks in modern AI research. Existing LLM-driven scientific discovery systems are typically limited to one-shot prompting on static literature snapshots and are validated only against contemporary judges such as human reviewers, agent peer review, wet-lab assays, or self-evaluation, leaving open whether they can anticipate future trends. We present Continuous Knowledge Metabolism (CKM), an AI workflow for hypothesis generation with three key capabilities: (i) continuous literature metabolism via sliding windows that maintain an evolving knowledge state; (ii) predictive evaluation, which grades hypotheses against papers published after the generation window; and (iii) practitioner-grade failure detection that diagnoses workflow failure modes from its outputs. On a 50-topic machine learning benchmark, CKM-Lite produces at least one validated hypothesis on 72% of topics (36 out of 50), more than doubling a one-shot baseline (30%) at approximately 3 dollars per topic and achieving 91% lower token cost. Validated hypotheses precede their matched papers by an average of 404 days (55 hits across 36 topics; median 399 days, range 66-757 days). Broadly, predictive validation against future literature provides a falsifiable, low-cost alternative to contemporary-judge evaluation protocols and can be applied wherever a corpus has dated publication records.

AI科研趋势预测假设生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。