让科研智能体自动进化科学原理,提升发现效率与泛化能力。
Principle-Evolvable Scientific Discovery via Uncertainty Minimization
- 通过贝叶斯优化动态扩展原理空间,实现科学假设的自主演化。
- 在4个基准上平均解质量达90.81%~93.15%,比现有方法提升29.7%~31.1%。
- 适合需要持续探索新理论的跨领域科研场景,兼容多种大模型底座。
基于大语言模型的科研智能体虽加速了科学发现,但常因固守初始先验而效率低下。现有方法多在静态假设空间中运行,限制了新现象的发现,导致基础理论失效时产生计算浪费。为此,我们提出将焦点从搜索假设转向演化科学原理。本文介绍PiEvo框架,将科学发现建模为对不断扩展的原理空间进行贝叶斯优化。通过融合基于高斯过程的信息导向假说选择与异常驱动的数据增强机制,PiEvo使智能体能自主完善其理论世界观。在四个基准上的评估表明:(1) 平均解质量达90.81%~93.15%,较当前最优水平提升29.7%~31.1%;(2) 通过优化紧凑的原理空间显著降低样本复杂度,收敛步数提速83.3%;(3) 在多样科学领域和不同LLM底座下均保持鲁棒性能。代码已公开于github.com/amair-lab/PiEvo。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based scientific agents have accelerated scientific discovery, yet they often suffer from significant inefficiencies due to adherence to fixed initial priors. Existing approaches predominantly operate within a static hypothesis space, which restricts the discovery of novel phenomena, resulting in computational waste when baseline theories fail. To address this, we propose shifting the focus from searching hypotheses to evolving the underlying scientific principles. We present PiEvo, a principle-evolvable framework that treats scientific discovery as Bayesian optimization over an expanding principle space. By integrating Information-Directed Hypothesis Selection via Gaussian Process and an anomaly-driven augmentation mechanism, PiEvo enables agents to autonomously refine their theoretical worldview. Evaluation across four benchmarks demonstrates that PiEvo (1) achieves an average solution quality of up to 90.81%~93.15%, representing a 29.7%~31.1% improvement over the state-of-the-art, (2) attains an 83.3% speedup in convergence step via significantly reduced sample complexity by optimizing the compact principle space, and (3) maintains robust performance across diverse scientific domains and LLM backbones. Code is publicly available at \hyperlink{https://github.com/amair-lab/PiEvo}{github.com/amair-lab/PiEvo}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。