arXiv:2510.08317physics.comp-phastro-ph.IM2025-10被引 4

用大模型生成有逻辑的数学表达式,让自动发现公式更准确可解释。

Iterated Agent for Symbolic Regression

  • 用大模型根据自然语言理由生成候选公式,引导搜索方向
  • 在费曼数据库上表现优于多个基线,抗噪能力强
  • 适合需要物理意义解释的科研场景,如高能物理参数建模

符号回归(SR)是通过数据自动发现数学表达式的基石,但常受搜索空间组合爆炸和过拟合困扰。传统基于遗传编程的方法在语法层面探索,往往产生过于复杂、难以理解的模型。本文提出IdeaSearchFitter框架,将大语言模型(LLMs)作为语义算子嵌入进化搜索中,通过自然语言推理生成候选表达式,使发现结果兼具准确性与概念一致性。我们在多种任务中验证其有效性:在费曼符号回归数据库(FSReD)上表现优异且抗噪;在真实世界数据中发现符合机理、精度与复杂度平衡良好的模型;并为高能物理前沿中的部分子分布函数推导出简洁、物理驱动的参数化形式。IdeaSearchFitter是公开框架IdeaSearch的专用模块,网址为https://www.ideasearch.cn/。

原文摘要 · Abstract (English)

Symbolic regression (SR), the automated discovery of mathematical expressions from data, is a cornerstone of scientific inquiry. However, it is often hindered by the combinatorial explosion of the search space and a tendency to overfit. Popular methods, rooted in genetic programming, explore this space syntactically, often yielding overly complex, uninterpretable models. This paper introduces IdeaSearchFitter, a framework that employs Large Language Models (LLMs) as semantic operators within an evolutionary search. By generating candidate expressions guided by natural-language rationales, our method biases discovery towards models that are not only accurate but also conceptually coherent and interpretable. We demonstrate IdeaSearchFitter's efficacy across diverse challenges: it achieves competitive, noise-robust performance on the Feynman Symbolic Regression Database (FSReD), outperforming several strong baselines; discovers mechanistically aligned models with good accuracy-complexity trade-offs on real-world data; and derives compact, physically-motivated parametrizations for Parton Distribution Functions in a frontier high-energy physics application. IdeaSearchFitter is a specialized module within our broader iterated agent framework, IdeaSearch, which is publicly available at https://www.ideasearch.cn/.

符号回归大模型科学发现高能物理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。