用科学解释框架让AI生成可验证的假设,超越单纯预测。
DN-Hypo-Pipeline: An AI-Driven Workflow for Generating Hypotheses using Large Language Models and Scientific Explanations
- 基于科学解释理论构建分层推理框架
- 生成的假设经专家和LLM评估显著更优
- 适合需要可解释性创新的研究者
现代人工智能擅长预测但无法解释。从大语言模型到科学领域AI系统,当前机器仅通过重组文献中已有模式回答‘是什么’,却难以基于底层原理推导‘为什么’,而解释才是科学发现的核心。本文探讨能否将科学解释结构转化为机器生成假设的指导原则。提出DN-Hypo-Pipeline框架,采用分层解释理论架构:赫梅尔的演绎-自然律(DN)模型提供假设形式与逻辑有效性,萨尔蒙的因果过程理论限定搜索范围,阿姆斯特朗关于定律是普遍关系的观点则连接现象构成过程与潜在规律。该框架不搜索已有文献,而是探索支配现象的原理:给定待解释现象,抽象其形成过程中涉及的普遍项,检索相关定律,并演绎重构新可检验解释。在数据科学建模中评估,由该框架生成的假设经大模型与人类专家评判均显著优于直接提示生成结果。关键成果:将两个得分最高的假设转化为新型算法——一个在仅轻微性能损失下降低Transformer理论复杂度,另一个以极少参数实现竞争性准确率。
原文摘要 · Abstract (English)
Modern artificial intelligence excels at prediction but cannot explain. From large language models to AI-for-science systems, today's machines answer what by recombining patterns already present in the human literature, yet they cannot reason out why a phenomenon must arise from underlying principles even though explanation, not prediction, lies at the heart of scientific discovery. Here we ask whether the structure of scientific explanation can be operationalized to guide how a machine generates hypotheses. We introduce DN-Hypo-Pipeline, a hypothesis-generation framework that adopts a layered, explanation-theoretic scaffold: Hempel's deductive-nomological (DN) model supplies the output form and deductive validity of a hypothesis, Salmon's causal-process account supplies an organizing constraint on where to search for the governing laws, and Armstrong's view of laws as relations between universals supplies the bridge from a phenomenon's constituent processes to the laws that may be associated with it. Rather than searching the space of what has been written, the framework searches the space of what principles govern a phenomenon: given an explanandum, it abstracts the universals instantiated in the phenomenon's formation process, retrieves the laws relating those universals, and deductively reconstructs a new, testable explanation. Evaluated in data-science modeling and judged by both LLMs and human experts, hypotheses generated through this principled reasoning significantly outperform those from direct prompting. Crucially, we translated the two highest-scoring hypotheses into novel algorithms one that reduces the Transformer's theoretical complexity with only minimal performance loss, and another that achieves competitive accuracy with substantially fewer parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。