用语言模型控制搜索路径,提升方程发现的效率与准确性
Language models guide symbolic equation discovery by controlling search
- 让语言模型指挥符号回归的搜索过程,而非直接选公式
- 在74个AI-Feynman任务中,准确率与复杂度平衡最优
- 适合需要高效探索科学规律的研究者
科学方程发现需融合广泛领域先验与严格数值验证。符号回归提供数值基础,但面临组合爆炸问题;许多语言模型系统直接让模型提出或选择公式。本文测试不同分工方式:语言模型作为方程作者、候选决策者或搜索控制器,对比端到端语言模型与纯数值基线。在提出的控制器设定中,语言模型指定变量、运算符、变换和搜索深度;符号回归枚举并拟合表达式;确定性指标决定保留。在74个AI-Feynman方程和七个复杂公式恢复任务中,搜索控制取得最佳精度、复杂度、稳定性和成本平衡。在独立电池数据集上,LLM-PySR成功识别出早期电压曲线偏移与循环寿命之间的紧凑分段线性关系。结果表明,语言模型应主导假设探索,而非决定公式留存。
原文摘要 · Abstract (English)
Scientific equation discovery must combine broad domain priors with strict numerical testing. Symbolic regression supplies numerical grounding but faces a combinatorial search space, whereas many language-model systems ask the model to propose or select formulas directly. We test a different division of labour. We compare role specifications in which the language model acts as equation author, candidate decider or search controller, alongside end-to-end language-model and purely numerical baselines. In the controller setting we propose here, implemented as LLM-PySR, language models specify variables, operators, transformations and search depth; symbolic regression enumerates and fits expressions; and deterministic metrics govern retention. Across 74 AI-Feynman equations and seven complex formula-recovery tasks, search control achieved the strongest observed balance of accuracy, complexity, stability and cost. On an independent battery dataset, LLM-PySR identified a compact piecewise-linear relation between early voltage-curve displacement and cycle life. The results suggest that language models should shape hypothesis exploration rather than decide which equations survive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。