arXiv:2505.04651cs.CLcs.LG2025-05被引 15

用大模型生成并验证科学假说,提升研究效率与可解释性。

Scientific Hypothesis Generation and Validation: Methods, Datasets, and Future Directions

  • 结合检索增强、知识图谱和因果推理,实现智能假说生成。
  • 对比传统符号系统与现代大模型,揭示性能与可解释性权衡。
  • 适合科研人员探索新方向,尤其关注跨学科发现的团队。

大型语言模型正在通过信息融合、潜在关系挖掘和推理增强,重塑科学假说的生成与验证方式。本文系统梳理了基于大模型的方法,包括符号框架、生成模型、混合系统和多智能体架构,分析了检索增强生成、知识图谱补全、模拟、因果推断及工具辅助推理等技术,强调其在可解释性、新颖性和领域适配性之间的权衡。文章对比了早期符号发现系统(如 BACON、KEKADA)与现代大模型流程,后者利用上下文学习、微调、检索和符号锚定实现领域适应。在验证方面,涵盖模拟、人机协作、因果建模和不确定性量化,强调开放世界中的迭代评估。调研覆盖生物医学、材料科学、环境科学和社会科学领域的数据集,引入 AHTech 与 CSKG-600 等新资源。最后提出发展路线图,聚焦新颖性感知生成、多模态-符号融合、人在回路系统及伦理保障,将大模型定位为可信赖、可扩展的科学发现代理。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are transforming scientific hypothesis generation and validation by enabling information synthesis, latent relationship discovery, and reasoning augmentation. This survey provides a structured overview of LLM-driven approaches, including symbolic frameworks, generative models, hybrid systems, and multi-agent architectures. We examine techniques such as retrieval-augmented generation, knowledge-graph completion, simulation, causal inference, and tool-assisted reasoning, highlighting trade-offs in interpretability, novelty, and domain alignment. We contrast early symbolic discovery systems (e.g., BACON, KEKADA) with modern LLM pipelines that leverage in-context learning and domain adaptation via fine-tuning, retrieval, and symbolic grounding. For validation, we review simulation, human-AI collaboration, causal modeling, and uncertainty quantification, emphasizing iterative assessment in open-world contexts. The survey maps datasets across biomedicine, materials science, environmental science, and social science, introducing new resources like AHTech and CSKG-600. Finally, we outline a roadmap emphasizing novelty-aware generation, multimodal-symbolic integration, human-in-the-loop systems, and ethical safeguards, positioning LLMs as agents for principled, scalable scientific discovery.

科学发现大模型假说生成综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。